{
  "id": 67376,
  "title": "Info for the challenge",
  "url": "/competitions/PLAsTiCC-2018/discussion/67376",
  "author_name": "Renee Hlozek",
  "post_date": "2018-10-02T00:41:02.823000",
  "votes": 28,
  "comment_count": 47,
  "views": 0,
  "content": "<p>Hi Kagglers</p>\n\n<p>We are so pleased that so many of you are already interested in the challenge, that mimics the kinds of data we are going to get from LSST. We have tried to give you as much information as you need for the challenge, and a bit of explanation about the astronomical concepts, but you don't need to be guided too strongly by any of that!&nbsp;</p>\n\n<p>We hope you are getting familiar with the data through the starter kit here:&nbsp;&nbsp;<a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a>\nIt contains demos of the data and classification algorithms, and gives a sense for what has been tried before.</p>\n\n<p>We also have shared the data&nbsp;note&nbsp;with the public and astronomers on the arXiv too: <a href=\"https://arxiv.org/abs/1810.00001\">https://arxiv.org/abs/1810.00001</a>\nwhere lots of astronomy papers are shared (it is a great place to learn about astronomy more generally).</p>\n\n<p>There has been lots of talk about metrics on the discussion boards. FYI we submitted a paper which talks a little bit about metrics for challenges like PLAsTiCC: <a href=\"https://arxiv.org/abs/1809.11145\">https://arxiv.org/abs/1809.11145</a> </p>\n\n<p>The  release note is geared for people without astronomy training, while the latter is more technical, and geared for astronomers/statistical enthusiasts.</p>\n\n<p>Hope you enjoy them both, and the challenge.</p>\n\n<ul>\n<li>Renée for the PLAsTiCC team</li>\n</ul>",
  "messages": [
    {
      "id": 397138,
      "postDate": "2018-10-02T00:41:02.823Z",
      "content": "<p>Hi Kagglers</p>\n\n<p>We are so pleased that so many of you are already interested in the challenge, that mimics the kinds of data we are going to get from LSST. We have tried to give you as much information as you need for the challenge, and a bit of explanation about the astronomical concepts, but you don't need to be guided too strongly by any of that!&nbsp;</p>\n\n<p>We hope you are getting familiar with the data through the starter kit here:&nbsp;&nbsp;<a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a>\nIt contains demos of the data and classification algorithms, and gives a sense for what has been tried before.</p>\n\n<p>We also have shared the data&nbsp;note&nbsp;with the public and astronomers on the arXiv too: <a href=\"https://arxiv.org/abs/1810.00001\">https://arxiv.org/abs/1810.00001</a>\nwhere lots of astronomy papers are shared (it is a great place to learn about astronomy more generally).</p>\n\n<p>There has been lots of talk about metrics on the discussion boards. FYI we submitted a paper which talks a little bit about metrics for challenges like PLAsTiCC: <a href=\"https://arxiv.org/abs/1809.11145\">https://arxiv.org/abs/1809.11145</a> </p>\n\n<p>The  release note is geared for people without astronomy training, while the latter is more technical, and geared for astronomers/statistical enthusiasts.</p>\n\n<p>Hope you enjoy them both, and the challenge.</p>\n\n<ul>\n<li>Renée for the PLAsTiCC team</li>\n</ul>",
      "rawMarkdown": "Hi Kagglers\n\nWe are so pleased that so many of you are already interested in the challenge, that mimics the kinds of data we are going to get from LSST. We have tried to give you as much information as you need for the challenge, and a bit of explanation about the astronomical concepts, but you don't need to be guided too strongly by any of that!&nbsp;\n\nWe hope you are getting familiar with the data through the starter kit here:&nbsp;&nbsp;https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\nIt contains demos of the data and classification algorithms, and gives a sense for what has been tried before.\n\nWe also have shared the data&nbsp;note&nbsp;with the public and astronomers on the arXiv too: https://arxiv.org/abs/1810.00001\nwhere lots of astronomy papers are shared (it is a great place to learn about astronomy more generally).\n\nThere has been lots of talk about metrics on the discussion boards. FYI we submitted a paper which talks a little bit about metrics for challenges like PLAsTiCC: https://arxiv.org/abs/1809.11145 \n\nThe  release note is geared for people without astronomy training, while the latter is more technical, and geared for astronomers/statistical enthusiasts.\n\nHope you enjoy them both, and the challenge.\n\n- Renée for the PLAsTiCC team",
      "votes": 28
    },
    {
      "id": 399104,
      "postDate": "2018-10-05T07:32:39.390Z",
      "content": "<p>Hi Renée,</p>\n\n<p>May I ask you a question:</p>\n\n<p>class_99 objects - can they actually be objects of different nature, the only thing which is in common for them is that we didn't see them in train set? This way I mean they are not a real class.</p>",
      "rawMarkdown": "Hi Renée,\n\nMay I ask you a question:\n\nclass_99 objects - can they actually be objects of different nature, the only thing which is in common for them is that we didn't see them in train set? This way I mean they are not a real class.",
      "votes": 3
    },
    {
      "id": 399261,
      "postDate": "2018-10-05T14:12:04.450Z",
      "content": "<p>hi Anna, what is means is that because LSST will be a new telescope which is better than ones we have at the moment, we can expect LSST to find objects that we haven't found before.  So we'll have no information on what the objects might be, including their class.</p>",
      "rawMarkdown": "hi Anna, what is means is that because LSST will be a new telescope which is better than ones we have at the moment, we can expect LSST to find objects that we haven't found before.  So we'll have no information on what the objects might be, including their class.",
      "votes": 4
    },
    {
      "id": 414554,
      "postDate": "2018-11-03T01:37:48.293Z",
      "content": "<p>Hi, Is it allowed to use pre-trained model's weights in this competition ?</p>",
      "rawMarkdown": "Hi, Is it allowed to use pre-trained model's weights in this competition ?",
      "votes": 2,
      "replies": [
        {
          "id": 414601,
          "postDate": "2018-11-03T06:05:50.840Z",
          "content": "<p>That would be very interesting indeed, thanks for asking.  I hope it is OK provided we share which pretrained model we use in a topic in the forum.</p>",
          "rawMarkdown": "That would be very interesting indeed, thanks for asking.  I hope it is OK provided we share which pretrained model we use in a topic in the forum."
        },
        {
          "id": 436759,
          "postDate": "2018-12-10T22:35:42.310Z",
          "content": "<p>I am using pre-trained ResNet50's weights:<br>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/tag/v0.2\"></a><a href=\"https://github.com/fchollet/deep-learning-models/releases/tag/v0.2\">https://github.com/fchollet/deep-learning-models/releases/tag/v0.2</a><br>\nI hope it is allowed as CPMP writes.</p>",
          "rawMarkdown": "I am using pre-trained ResNet50's weights:<br>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/tag/v0.2\">https://github.com/fchollet/deep-learning-models/releases/tag/v0.2</a><br>\nI hope it is allowed as CPMP writes."
        },
        {
          "id": 438049,
          "postDate": "2018-12-13T02:47:29.967Z",
          "content": "<p>@Mickey, how were you able to deal with uneven samples for training the cnn?</p>",
          "rawMarkdown": "@Mickey, how were you able to deal with uneven samples for training the cnn?"
        },
        {
          "id": 438093,
          "postDate": "2018-12-13T04:49:11.890Z",
          "content": "<p>Hi dylonLL,<br>\nI used linear interpolation plus some noise.<br>\nI tried Gaussian Process Regression and Spline interpolation, but I could not get any good results.</p>",
          "rawMarkdown": "Hi dylonLL,<br>\nI used linear interpolation plus some noise.<br>\nI tried Gaussian Process Regression and Spline interpolation, but I could not get any good results.",
          "votes": 1
        },
        {
          "id": 438095,
          "postDate": "2018-12-13T04:55:21.523Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks"
        }
      ]
    },
    {
      "id": 416974,
      "postDate": "2018-11-07T14:56:41.070Z",
      "content": "<p>Is it possible to  upload the data from this git repo on Kaggle? There is no license file:</p>\n\n<p><a href=\"https://github.com/lsst/throughputs/tree/master/baseline\">https://github.com/lsst/throughputs/tree/master/baseline</a></p>\n\n<p>I am asking because I'd like to explain how I managed to reproduce the throughput curves using only pandas and matplotlib.  It may help participants understand what is measured.</p>\n\n<p>My throughput curves:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416974/10631/throughput.png\" alt=\"throughput curves\"></p>",
      "rawMarkdown": "Is it possible to  upload the data from this git repo on Kaggle? There is no license file:\n\nhttps://github.com/lsst/throughputs/tree/master/baseline\n\nI am asking because I'd like to explain how I managed to reproduce the throughput curves using only pandas and matplotlib.  It may help participants understand what is measured.\n\nMy throughput curves:\n\n![throughput curves][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416974/10631/throughput.png",
      "votes": 1,
      "replies": [
        {
          "id": 418548,
          "postDate": "2018-11-10T05:26:23.567Z",
          "content": "<p>I want to know if we can use these data, too. If it's ok, we can make analysis more rich in variety. </p>",
          "rawMarkdown": "I want to know if we can use these data, too. If it's ok, we can make analysis more rich in variety. "
        },
        {
          "id": 431267,
          "postDate": "2018-12-02T00:46:57.467Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 439515,
      "postDate": "2018-12-15T17:17:50.403Z",
      "content": "<p><a href=\"/kyleboone\">@kyleboone</a> <a href=\"/gsnarayan\">@gsnarayan</a></p>\n\n<p>To clarify: are the given data corrected for the filters opacity? Yes/No ? </p>",
      "rawMarkdown": "@kyleboone @gsnarayan\n\nTo clarify: are the given data corrected for the filters opacity? Yes/No ? ",
      "replies": [
        {
          "id": 439533,
          "postDate": "2018-12-15T18:23:51.340Z",
          "content": "<p>As far as I know, yes, they are.</p>",
          "rawMarkdown": "As far as I know, yes, they are.",
          "votes": 1
        },
        {
          "id": 439538,
          "postDate": "2018-12-15T18:39:13.337Z",
          "content": "<p>thank you, the confirmation from org would be also nice. i also assumed filters opacity is already taken into account for data for my features, but did not notice the explicit formulation of this in data note</p>",
          "rawMarkdown": "thank you, the confirmation from org would be also nice. i also assumed filters opacity is already taken into account for data for my features, but did not notice the explicit formulation of this in data note"
        }
      ]
    },
    {
      "id": 436696,
      "postDate": "2018-12-10T19:00:10.523Z",
      "content": "<p>I was unable to unpack test_set.csv.zip on two different computers, one of them quite new,  with multiple download attempts. The issue appears to be the size of the file, since I can unpack everything else. I there a way to download it in multiple smaller subsets?</p>",
      "rawMarkdown": "I was unable to unpack test_set.csv.zip on two different computers, one of them quite new,  with multiple download attempts. The issue appears to be the size of the file, since I can unpack everything else. I there a way to download it in multiple smaller subsets?"
    },
    {
      "id": 418562,
      "postDate": "2018-11-10T05:48:31.790Z",
      "content": "<p>I have a question about flux values. \nAs mentioned above by CPMP, each filters have specific throughput characteristics. Are flux values in the competitions data already corrected from these effect, or raw values from censor?\nIn the latter case, I think we have to correct those to able to equaly treat each passband. </p>",
      "rawMarkdown": "I have a question about flux values. \nAs mentioned above by CPMP, each filters have specific throughput characteristics. Are flux values in the competitions data already corrected from these effect, or raw values from censor?\nIn the latter case, I think we have to correct those to able to equaly treat each passband. ",
      "replies": [
        {
          "id": 420055,
          "postDate": "2018-11-13T01:42:00.600Z",
          "content": "<p>The flux values are already corrected for the difference in sensitivity. We've worked to make sure that there is no specialized domain knowledge you need to take part in this competition, and we'd consider this to be one such example. Note that for a source with constant brightness across all passbands, the errors will still be higher for passbands with lower sensitivity. </p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "The flux values are already corrected for the difference in sensitivity. We've worked to make sure that there is no specialized domain knowledge you need to take part in this competition, and we'd consider this to be one such example. Note that for a source with constant brightness across all passbands, the errors will still be higher for passbands with lower sensitivity. \n\nCheers,\n\n-Gautham for the PLAsTiCC team\n\n",
          "votes": 3
        },
        {
          "id": 421044,
          "postDate": "2018-11-14T14:07:10.637Z",
          "content": "<p>Appreciate the information.  I have another question on flux. </p>\n\n<p>What is the unit of flux provided in the data? I suppose that it's in the unit of [erg s-1 cm-2 Hz-1] and wanted to confirm whether my understanding is correct. I think this may help to understand the nature of transients.</p>",
          "rawMarkdown": "Appreciate the information.  I have another question on flux. \n\nWhat is the unit of flux provided in the data? I suppose that it's in the unit of [erg s-1 cm-2 Hz-1] and wanted to confirm whether my understanding is correct. I think this may help to understand the nature of transients.",
          "votes": 1
        },
        {
          "id": 421072,
          "postDate": "2018-11-14T14:43:46.140Z",
          "content": "<p>Hi Jun,</p>\n\n<p>No, the fluxes are in normalized counts, and we've not provided a zeropoint to convert them into an absolute magnitude or some calibrated units like janskys or [erg s-1 cm-2 Hz-1]. </p>\n\n<p>It was a deliberate choice to exclude the normalization - models that are built to classify the data that are sensitive to the absolute flux will exhibit a bias with increasing redshift and with increasingly dusty environments. There's no way to avoid this bias as it's sort of intrinsic to the way the Universe works - as we go to higher distances or redshifts, we only see the most bright objects, so we're not sampling the entire population. This also holds as the amount of dust along the line of sight increases. We intend to use the models developed for this challenge in actual science pipelines, so any bias in the models propagates into any cosmological bias with LSST.  This would render the model unusable for many of the science applications we care about. This defeats the purpose of the challenge.  </p>\n\n<p>We know it'd be nice to have an absolute brightness or color information, and indeed it is necessary if we want to learn about the nature of transients. That said, we also know the consequences of giving that information out on our analysis downstream from classification, and the problems that we will introduce by using a model that is trained on a feature that is guaranteed to exhibit bias. We want people to use the shape of the light curve, and flux ratios as much as possible, rather than the overall normalization.</p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "Hi Jun,\n\nNo, the fluxes are in normalized counts, and we've not provided a zeropoint to convert them into an absolute magnitude or some calibrated units like janskys or [erg s-1 cm-2 Hz-1]. \n\nIt was a deliberate choice to exclude the normalization - models that are built to classify the data that are sensitive to the absolute flux will exhibit a bias with increasing redshift and with increasingly dusty environments. There's no way to avoid this bias as it's sort of intrinsic to the way the Universe works - as we go to higher distances or redshifts, we only see the most bright objects, so we're not sampling the entire population. This also holds as the amount of dust along the line of sight increases. We intend to use the models developed for this challenge in actual science pipelines, so any bias in the models propagates into any cosmological bias with LSST.  This would render the model unusable for many of the science applications we care about. This defeats the purpose of the challenge.  \n\nWe know it'd be nice to have an absolute brightness or color information, and indeed it is necessary if we want to learn about the nature of transients. That said, we also know the consequences of giving that information out on our analysis downstream from classification, and the problems that we will introduce by using a model that is trained on a feature that is guaranteed to exhibit bias. We want people to use the shape of the light curve, and flux ratios as much as possible, rather than the overall normalization.\n\nCheers,\n\n-Gautham for the PLAsTiCC team",
          "votes": 5
        },
        {
          "id": 421089,
          "postDate": "2018-11-14T15:09:16.270Z",
          "content": "<p>Hi Gautham,</p>\n\n<p>Thank you very much for the detailed explanation. \nAs you mentioned, to exploit rest-flame information, such as brightness or color, were what came in my mind at first, but I understand your effort not to introduce biases into the model and I completely agree with your comments. I'll keep trying to utilize light curves as possible. </p>\n\n<p>Appreciate the quick response.\nCheers,\nJun</p>",
          "rawMarkdown": "Hi Gautham,\n\nThank you very much for the detailed explanation. \nAs you mentioned, to exploit rest-flame information, such as brightness or color, were what came in my mind at first, but I understand your effort not to introduce biases into the model and I completely agree with your comments. I'll keep trying to utilize light curves as possible. \n\nAppreciate the quick response.\nCheers,\nJun",
          "votes": 1
        },
        {
          "id": 421127,
          "postDate": "2018-11-14T16:08:29.153Z",
          "content": "<p>Hi Gautham,</p>\n\n<p>Thanks for kindly explanation. \nDoes \"normalized counts\" means that flux values are also corrected for differences of passband width and so we can directly compare values of whole passbands without any conversion?</p>",
          "rawMarkdown": "Hi Gautham,\n\nThanks for kindly explanation. \nDoes \"normalized counts\" means that flux values are also corrected for differences of passband width and so we can directly compare values of whole passbands without any conversion?"
        },
        {
          "id": 421143,
          "postDate": "2018-11-14T16:30:28.620Z",
          "content": "<p>Hi Takuya, </p>\n\n<p>The correction for different passband widths is folded into the correction for passband sensitivity differences, so yes. Basically, you can use the counts provided, and again, we don't expect people to make any special conversion because that sort of knowledge would give astronomers an edge in this contest. We've tried to make sure things are as balanced as possible between making the data easy to use, and providing background through the data note and the starter kit.</p>\n\n<p>Best,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "Hi Takuya, \n\nThe correction for different passband widths is folded into the correction for passband sensitivity differences, so yes. Basically, you can use the counts provided, and again, we don't expect people to make any special conversion because that sort of knowledge would give astronomers an edge in this contest. We've tried to make sure things are as balanced as possible between making the data easy to use, and providing background through the data note and the starter kit.\n\nBest,\n\n-Gautham for the PLAsTiCC team",
          "votes": 2
        },
        {
          "id": 421179,
          "postDate": "2018-11-14T17:11:27.317Z",
          "content": "<p>Thanks, again. \nYeah, I know your intention to expect for us to compete regardless of domain knowledge . However in any competition, I think it is the most important thing to understand the data more deeply. \nSorry for insistant questions, but those data clarifications are necessary for our strategy. I guess I'll question you sometime again, but thank you.\nTakuya</p>",
          "rawMarkdown": "Thanks, again. \nYeah, I know your intention to expect for us to compete regardless of domain knowledge . However in any competition, I think it is the most important thing to understand the data more deeply. \nSorry for insistant questions, but those data clarifications are necessary for our strategy. I guess I'll question you sometime again, but thank you.\nTakuya"
        }
      ]
    },
    {
      "id": 408520,
      "postDate": "2018-10-23T03:03:58.080Z",
      "content": "<p>Hi, this looks like a quite challenging competition, thanks for hosting it!  </p>\n\n<p>I have a question about a part of the starter kit kernel.  In section 1.b it reads:</p>\n\n<blockquote>\n  <p>We'll give you a few external resources for these events in a companion Kernel, if you are determined to augment the training set. </p>\n</blockquote>\n\n<p>Has this companion kernel been released?</p>",
      "rawMarkdown": "Hi, this looks like a quite challenging competition, thanks for hosting it!  \n\nI have a question about a part of the starter kit kernel.  In section 1.b it reads:\n\n&gt; We'll give you a few external resources for these events in a companion Kernel, if you are determined to augment the training set. \n\nHas this companion kernel been released?",
      "replies": [
        {
          "id": 409018,
          "postDate": "2018-10-23T18:12:26.357Z",
          "content": "<p>Hi @CPMP,</p>\n\n<p>It has - it went live with the challenge. I expect it'll get noticed more as people get further into the challenge. The resources are listed in a table at the bottom.</p>\n\n<p><a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-classification-demo\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-classification-demo</a></p>\n\n<p>Best,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "Hi @CPMP,\n\nIt has - it went live with the challenge. I expect it'll get noticed more as people get further into the challenge. The resources are listed in a table at the bottom.\n\nhttps://www.kaggle.com/michaelapers/the-plasticc-astronomy-classification-demo\n\nBest,\n\n-Gautham for the PLAsTiCC team",
          "votes": 1
        },
        {
          "id": 409029,
          "postDate": "2018-10-23T18:29:59.060Z",
          "content": "<p>Thanks a lot, look there is a LOT to explore!</p>",
          "rawMarkdown": "Thanks a lot, look there is a LOT to explore!"
        }
      ]
    },
    {
      "id": 407986,
      "postDate": "2018-10-22T05:25:02.087Z",
      "content": "<p>Hello Renée and PLAsTiCC team, thank you for holding such an interesting competition.<br>\nbtw, I have a few questions about <code>detected</code> and <code>flux_err, hostgal_photoz_err</code>in the data set. <br><br>\nI found many samples, which satisfy <code>detected == 1</code> &amp; <code>flux &lt; 0</code> at the same time. <br> \nI think it's strange because <code>data_note.pdf</code> says <br>\n<code>If detected = 1, the object's brightness is significantly different at the 3\\sigma level relative to the reference template. This is given as Boolean flag.</code> <br> <br></p>\n\n<p>In addition, I couldn't understand how <code>flux_err, hostgal_photoz_err</code> are caluclated, which may have something to do with <code>detected</code>. <br><br>\nMy Question: <br>\n1. How is <code>flux_err</code> and <code>hostgal_photoz_err</code> caluclated? <br>\n2. How is <code>detected</code> calculated ?</p>",
      "rawMarkdown": "Hello Renée and PLAsTiCC team, thank you for holding such an interesting competition.<br>\nbtw, I have a few questions about `detected` and `flux_err, hostgal_photoz_err`in the data set. <br><br>\nI found many samples, which satisfy `detected == 1` &amp; `flux &lt; 0` at the same time. <br> \nI think it's strange because `data_note.pdf` says <br>\n`If detected = 1, the object's brightness is significantly different at the 3\\sigma level relative to the reference template. This is given as Boolean flag.` <br> <br>\n\nIn addition, I couldn't understand how `flux_err, hostgal_photoz_err` are caluclated, which may have something to do with `detected`. <br><br>\nMy Question: <br>\n1. How is `flux_err` and `hostgal_photoz_err` caluclated? <br>\n2. How is `detected` calculated ?",
      "replies": [
        {
          "id": 413861,
          "postDate": "2018-11-01T16:57:24.877Z",
          "content": "<p>I'm wondering if PLAsTiCC team have had chance to look at the comment above. thanks.</p>",
          "rawMarkdown": "I'm wondering if PLAsTiCC team have had chance to look at the comment above. thanks."
        },
        {
          "id": 413868,
          "postDate": "2018-11-01T17:06:30.680Z",
          "content": "<blockquote>\n  <p>detected == 1 &amp; flux &lt; 0 at the same time. </p>\n</blockquote>\n\n<p>This can be true is flux is so low that it is more than 3 std below the reference level.  That's how I understand this anyway, a confirmation would be great.</p>",
          "rawMarkdown": "&gt; detected == 1 &amp; flux &lt; 0 at the same time. \n\nThis can be true is flux is so low that it is more than 3 std below the reference level.  That's how I understand this anyway, a confirmation would be great.",
          "votes": 1
        },
        {
          "id": 413876,
          "postDate": "2018-11-01T17:32:03.130Z",
          "content": "<p>Hi CPMP, thanks for the reply. <br>\nI first thought same as you, but I found some negative flux samples which is not 3 std below the reference level. \nI also found some positive flux samples which is not 3 std above the reference level. please see the images.\n<br>\n<a href=\"https://imgur.com/srPyhuC\">https://imgur.com/srPyhuC</a> <br>\n<a href=\"https://imgur.com/vW7amPE\">https://imgur.com/vW7amPE</a></p>",
          "rawMarkdown": "Hi CPMP, thanks for the reply. <br>\nI first thought same as you, but I found some negative flux samples which is not 3 std below the reference level. \nI also found some positive flux samples which is not 3 std above the reference level. please see the images.\n<br>\nhttps://imgur.com/srPyhuC <br>\nhttps://imgur.com/vW7amPE",
          "votes": 2
        },
        {
          "id": 414045,
          "postDate": "2018-11-02T02:16:45.833Z",
          "content": "<p>my understanding is that: imagine you have a blinking star with average brightness above 3 sigma from the reference level. But at some moments it totally fades and the flux drops below reference levels, also by 3 sigma or more...  </p>",
          "rawMarkdown": "my understanding is that: imagine you have a blinking star with average brightness above 3 sigma from the reference level. But at some moments it totally fades and the flux drops below reference levels, also by 3 sigma or more...  ",
          "votes": 2
        },
        {
          "id": 414389,
          "postDate": "2018-11-02T16:37:17.483Z",
          "content": "<p>hmm, so the <code>detected</code> is dependent on the other rows? Anyway, we need to wait for the Host's comment for this problem.</p>",
          "rawMarkdown": "hmm, so the `detected` is dependent on the other rows? Anyway, we need to wait for the Host's comment for this problem."
        },
        {
          "id": 414462,
          "postDate": "2018-11-02T20:09:55.043Z",
          "content": "<p>yes, I am not an astronomer, so Host comments on this would be much better. It would be also nice to know how the flux is normalized, is the \"reference template\" the same for all data in train and the same as for the test. A bit more details on how it's done</p>",
          "rawMarkdown": "yes, I am not an astronomer, so Host comments on this would be much better. It would be also nice to know how the flux is normalized, is the \"reference template\" the same for all data in train and the same as for the test. A bit more details on how it's done"
        },
        {
          "id": 415280,
          "postDate": "2018-11-04T20:02:22.260Z",
          "content": "<p>I guess flux_err were calculated using error images/frames [poisson noise (==sqrt(flux/gain)) + readnoise], and detected == 1 means brightness is &gt; 3 std above the background (source) in the differential image (Img - Img_ref). Detected == 1 doesn't mean flux &gt; flux_err×3, it means the diff flux is at 3 std level of the background noise.</p>",
          "rawMarkdown": "I guess flux_err were calculated using error images/frames [poisson noise (==sqrt(flux/gain)) + readnoise], and detected == 1 means brightness is &gt; 3 std above the background (source) in the differential image (Img - Img_ref). Detected == 1 doesn't mean flux &gt; flux_err×3, it means the diff flux is at 3 std level of the background noise.",
          "votes": 2
        },
        {
          "id": 415907,
          "postDate": "2018-11-05T22:19:04.517Z",
          "content": "<blockquote>\n  <p><strong>mamasinkgs wrote</strong></p>\n  \n  <p>I found many samples, which satisfy <code>detected == 1</code> &amp; <code>flux &amp;lt; 0</code> at the same time. <br> </p>\n</blockquote>\n\n<p>The detection logic is based on <code>abs(SNR)</code>, so negative fluxes can result in a detection.</p>\n\n<blockquote>\n  <ol>\n  <li>How is <code>flux_err</code> and <code>hostgal_photoz_err</code> caluclated? <br></li>\n  </ol>\n</blockquote>\n\n<p>The reported <code>flux_err</code> is a statistical uncertainty from the number of (simulated) photo-electrons read from the CCDs, added in quadrature with the uncertainty arising from correcting the <code>flux</code> for the Milky Way reddening <code>mwebv</code> We aply the Milky Way reddening correction because it's likely the first pre-processing step that astronomers will use, and we felt that this was specialized domain knowledge that contestants without an astronomy background cannot be expected to know. </p>\n\n<p>Source detection algorithms are more complex. The <code>detected</code> flag is set with respect to efficiency curves that are determined by injecting fake sources with known brightness into images, and determining the efficiency of recovering these sources. The flux level at which to set <code>detected = 1</code> is determined from these curves, such that on average sources with S/N &gt;= 3 are treated as real. So <code>detected = 1</code> reflects the the image properties, not just the properties of a single source, and there may be small statistical fluctuations for any given observation. </p>\n\n<p><code>hostgal_photoz_err</code> is much more complicated (there's numerous techniques to determine the photo-z, some of which are ML-based) and we'll discuss this in papers after the end of the challenge. Photometric redshifts might be an interesting subject for a challenge in its own right. </p>\n\n<p>Best,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "&gt; **mamasinkgs wrote**\n\n&gt; I found many samples, which satisfy `detected == 1` &amp; `flux &lt; 0` at the same time. <br> \n\nThe detection logic is based on `abs(SNR)`, so negative fluxes can result in a detection.\n\n&gt; 1. How is `flux_err` and `hostgal_photoz_err` caluclated? <br>\n\nThe reported `flux_err` is a statistical uncertainty from the number of (simulated) photo-electrons read from the CCDs, added in quadrature with the uncertainty arising from correcting the `flux` for the Milky Way reddening `mwebv` We aply the Milky Way reddening correction because it's likely the first pre-processing step that astronomers will use, and we felt that this was specialized domain knowledge that contestants without an astronomy background cannot be expected to know. \n\nSource detection algorithms are more complex. The `detected` flag is set with respect to efficiency curves that are determined by injecting fake sources with known brightness into images, and determining the efficiency of recovering these sources. The flux level at which to set `detected = 1` is determined from these curves, such that on average sources with S/N &gt;= 3 are treated as real. So `detected = 1` reflects the the image properties, not just the properties of a single source, and there may be small statistical fluctuations for any given observation. \n\n`hostgal_photoz_err` is much more complicated (there's numerous techniques to determine the photo-z, some of which are ML-based) and we'll discuss this in papers after the end of the challenge. Photometric redshifts might be an interesting subject for a challenge in its own right. \n\nBest,\n\n-Gautham for the PLAsTiCC team",
          "votes": 2
        },
        {
          "id": 416033,
          "postDate": "2018-11-06T04:07:33.370Z",
          "content": "<p>Thanks for the explanations.</p>\n\n<blockquote>\n  <p>a statistical uncertainty </p>\n</blockquote>\n\n<p>Is it the standard deviation of the error, twice the std, something else?</p>",
          "rawMarkdown": "Thanks for the explanations.\n\n&gt; a statistical uncertainty \n\nIs it the standard deviation of the error, twice the std, something else?"
        },
        {
          "id": 416183,
          "postDate": "2018-11-06T10:50:46.863Z",
          "content": "<p>Thanks for the explanations !</p>",
          "rawMarkdown": "Thanks for the explanations !"
        },
        {
          "id": 416481,
          "postDate": "2018-11-06T17:47:48.560Z",
          "content": "<p>Hi @CPMP,</p>\n\n<blockquote>\n  <p>Is it the standard deviation of the error, twice the std, something else?</p>\n</blockquote>\n\n<p>One standard deviation for <code>flux_err</code></p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "Hi @CPMP,\n\n&gt; Is it the standard deviation of the error, twice the std, something else?\n\nOne standard deviation for `flux_err`\n\nCheers,\n\n-Gautham for the PLAsTiCC team",
          "votes": 2
        },
        {
          "id": 416488,
          "postDate": "2018-11-06T18:14:05.243Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        },
        {
          "id": 416597,
          "postDate": "2018-11-07T00:01:14.567Z",
          "content": "<p>Thank you for the update! <code>The flux level at which to set detected = 1 is determined from these curves, such that on average sources with S/N &gt;= 3 are treated as real. So detected = 1 reflects the image properties</code> -- does it mean that when detected = 0 from images, such points are usually disregarded by astronomers ?  </p>",
          "rawMarkdown": "Thank you for the update! ```The flux level at which to set detected = 1 is determined from these curves, such that on average sources with S/N &gt;= 3 are treated as real. So detected = 1 reflects the image properties``` -- does it mean that when detected = 0 from images, such points are usually disregarded by astronomers ?  \n"
        },
        {
          "id": 416668,
          "postDate": "2018-11-07T03:54:16.990Z",
          "content": "<p>Hi Blonde,</p>\n\n<blockquote>\n  <p>does it mean that when detected = 0 from images, such points are usually disregarded by astronomers ? </p>\n</blockquote>\n\n<p>There's lots of different approaches - some groups disregard these \"non-detections\",  while others do not. </p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "Hi Blonde,\n\n&gt;  does it mean that when detected = 0 from images, such points are usually disregarded by astronomers ? \n\nThere's lots of different approaches - some groups disregard these \"non-detections\",  while others do not. \n\nCheers,\n\n-Gautham for the PLAsTiCC team\n",
          "votes": 1
        },
        {
          "id": 417040,
          "postDate": "2018-11-07T16:36:21.333Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 422286,
      "postDate": "2018-11-16T02:39:44.777Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 422333,
          "postDate": "2018-11-16T04:51:18.527Z",
          "content": "<p>No it is not log scale, I asked the same question here: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69432#latest-409324\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69432#latest-409324</a></p>",
          "rawMarkdown": "No it is not log scale, I asked the same question here: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69432#latest-409324"
        },
        {
          "id": 422338,
          "postDate": "2018-11-16T04:56:12.623Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 423002,
          "postDate": "2018-11-17T08:34:13.067Z",
          "content": "<p>Desktop.  I fried my macbook pro with kaggle competitions, I'm no longer using it for long, cpu intensive runs.</p>",
          "rawMarkdown": "Desktop.  I fried my macbook pro with kaggle competitions, I'm no longer using it for long, cpu intensive runs."
        }
      ]
    },
    {
      "id": 404992,
      "postDate": "2018-10-16T17:52:34.700Z",
      "content": "<p>Thanks for links.</p>",
      "rawMarkdown": "Thanks for links."
    }
  ],
  "comments": [
    {
      "id": 399104,
      "author_name": "Anna Novikova",
      "author_url": "",
      "post_date": "2018-10-05T07:32:39.390000",
      "content": "<p>Hi Renée,</p>\n\n<p>May I ask you a question:</p>\n\n<p>class_99 objects - can they actually be objects of different nature, the only thing which is in common for them is that we didn't see them in train set? This way I mean they are not a real class.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 399261,
      "author_name": "Renee Hlozek",
      "author_url": "",
      "post_date": "2018-10-05T14:12:04.450000",
      "content": "<p>hi Anna, what is means is that because LSST will be a new telescope which is better than ones we have at the moment, we can expect LSST to find objects that we haven't found before.  So we'll have no information on what the objects might be, including their class.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 414554,
      "author_name": "Mickey",
      "author_url": "",
      "post_date": "2018-11-03T01:37:48.293000",
      "content": "<p>Hi, Is it allowed to use pre-trained model's weights in this competition ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 414601,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-03T06:05:50.840000",
          "content": "<p>That would be very interesting indeed, thanks for asking.  I hope it is OK provided we share which pretrained model we use in a topic in the forum.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 436759,
          "author_name": "Mickey",
          "author_url": "",
          "post_date": "2018-12-10T22:35:42.310000",
          "content": "<p>I am using pre-trained ResNet50's weights:<br>\n<a href=\"https://github.com/fchollet/deep-learning-models/releases/tag/v0.2\"></a><a href=\"https://github.com/fchollet/deep-learning-models/releases/tag/v0.2\">https://github.com/fchollet/deep-learning-models/releases/tag/v0.2</a><br>\nI hope it is allowed as CPMP writes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438049,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-12-13T02:47:29.967000",
          "content": "<p>@Mickey, how were you able to deal with uneven samples for training the cnn?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 438093,
          "author_name": "Mickey",
          "author_url": "",
          "post_date": "2018-12-13T04:49:11.890000",
          "content": "<p>Hi dylonLL,<br>\nI used linear interpolation plus some noise.<br>\nI tried Gaussian Process Regression and Spline interpolation, but I could not get any good results.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 438095,
          "author_name": "dylonLL",
          "author_url": "",
          "post_date": "2018-12-13T04:55:21.523000",
          "content": "<p>thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416974,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-11-07T14:56:41.070000",
      "content": "<p>Is it possible to  upload the data from this git repo on Kaggle? There is no license file:</p>\n\n<p><a href=\"https://github.com/lsst/throughputs/tree/master/baseline\">https://github.com/lsst/throughputs/tree/master/baseline</a></p>\n\n<p>I am asking because I'd like to explain how I managed to reproduce the throughput curves using only pandas and matplotlib.  It may help participants understand what is measured.</p>\n\n<p>My throughput curves:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416974/10631/throughput.png\" alt=\"throughput curves\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 418548,
          "author_name": "akiyama",
          "author_url": "",
          "post_date": "2018-11-10T05:26:23.567000",
          "content": "<p>I want to know if we can use these data, too. If it's ok, we can make analysis more rich in variety. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 431267,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-02T00:46:57.467000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 439515,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2018-12-15T17:17:50.403000",
      "content": "<p><a href=\"/kyleboone\">@kyleboone</a> <a href=\"/gsnarayan\">@gsnarayan</a></p>\n\n<p>To clarify: are the given data corrected for the filters opacity? Yes/No ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 439533,
          "author_name": "Kyle Boone",
          "author_url": "",
          "post_date": "2018-12-15T18:23:51.340000",
          "content": "<p>As far as I know, yes, they are.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 439538,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-12-15T18:39:13.337000",
          "content": "<p>thank you, the confirmation from org would be also nice. i also assumed filters opacity is already taken into account for data for my features, but did not notice the explicit formulation of this in data note</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 436696,
      "author_name": "Michael Walker",
      "author_url": "",
      "post_date": "2018-12-10T19:00:10.523000",
      "content": "<p>I was unable to unpack test_set.csv.zip on two different computers, one of them quite new,  with multiple download attempts. The issue appears to be the size of the file, since I can unpack everything else. I there a way to download it in multiple smaller subsets?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418562,
      "author_name": "akiyama",
      "author_url": "",
      "post_date": "2018-11-10T05:48:31.790000",
      "content": "<p>I have a question about flux values. \nAs mentioned above by CPMP, each filters have specific throughput characteristics. Are flux values in the competitions data already corrected from these effect, or raw values from censor?\nIn the latter case, I think we have to correct those to able to equaly treat each passband. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 420055,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-13T01:42:00.600000",
          "content": "<p>The flux values are already corrected for the difference in sensitivity. We've worked to make sure that there is no specialized domain knowledge you need to take part in this competition, and we'd consider this to be one such example. Note that for a source with constant brightness across all passbands, the errors will still be higher for passbands with lower sensitivity. </p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 421044,
          "author_name": "Jun Ernesto Okumura",
          "author_url": "",
          "post_date": "2018-11-14T14:07:10.637000",
          "content": "<p>Appreciate the information.  I have another question on flux. </p>\n\n<p>What is the unit of flux provided in the data? I suppose that it's in the unit of [erg s-1 cm-2 Hz-1] and wanted to confirm whether my understanding is correct. I think this may help to understand the nature of transients.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421072,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-14T14:43:46.140000",
          "content": "<p>Hi Jun,</p>\n\n<p>No, the fluxes are in normalized counts, and we've not provided a zeropoint to convert them into an absolute magnitude or some calibrated units like janskys or [erg s-1 cm-2 Hz-1]. </p>\n\n<p>It was a deliberate choice to exclude the normalization - models that are built to classify the data that are sensitive to the absolute flux will exhibit a bias with increasing redshift and with increasingly dusty environments. There's no way to avoid this bias as it's sort of intrinsic to the way the Universe works - as we go to higher distances or redshifts, we only see the most bright objects, so we're not sampling the entire population. This also holds as the amount of dust along the line of sight increases. We intend to use the models developed for this challenge in actual science pipelines, so any bias in the models propagates into any cosmological bias with LSST.  This would render the model unusable for many of the science applications we care about. This defeats the purpose of the challenge.  </p>\n\n<p>We know it'd be nice to have an absolute brightness or color information, and indeed it is necessary if we want to learn about the nature of transients. That said, we also know the consequences of giving that information out on our analysis downstream from classification, and the problems that we will introduce by using a model that is trained on a feature that is guaranteed to exhibit bias. We want people to use the shape of the light curve, and flux ratios as much as possible, rather than the overall normalization.</p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 421089,
          "author_name": "Jun Ernesto Okumura",
          "author_url": "",
          "post_date": "2018-11-14T15:09:16.270000",
          "content": "<p>Hi Gautham,</p>\n\n<p>Thank you very much for the detailed explanation. \nAs you mentioned, to exploit rest-flame information, such as brightness or color, were what came in my mind at first, but I understand your effort not to introduce biases into the model and I completely agree with your comments. I'll keep trying to utilize light curves as possible. </p>\n\n<p>Appreciate the quick response.\nCheers,\nJun</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421127,
          "author_name": "akiyama",
          "author_url": "",
          "post_date": "2018-11-14T16:08:29.153000",
          "content": "<p>Hi Gautham,</p>\n\n<p>Thanks for kindly explanation. \nDoes \"normalized counts\" means that flux values are also corrected for differences of passband width and so we can directly compare values of whole passbands without any conversion?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 421143,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-14T16:30:28.620000",
          "content": "<p>Hi Takuya, </p>\n\n<p>The correction for different passband widths is folded into the correction for passband sensitivity differences, so yes. Basically, you can use the counts provided, and again, we don't expect people to make any special conversion because that sort of knowledge would give astronomers an edge in this contest. We've tried to make sure things are as balanced as possible between making the data easy to use, and providing background through the data note and the starter kit.</p>\n\n<p>Best,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 421179,
          "author_name": "akiyama",
          "author_url": "",
          "post_date": "2018-11-14T17:11:27.317000",
          "content": "<p>Thanks, again. \nYeah, I know your intention to expect for us to compete regardless of domain knowledge . However in any competition, I think it is the most important thing to understand the data more deeply. \nSorry for insistant questions, but those data clarifications are necessary for our strategy. I guess I'll question you sometime again, but thank you.\nTakuya</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 408520,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-10-23T03:03:58.080000",
      "content": "<p>Hi, this looks like a quite challenging competition, thanks for hosting it!  </p>\n\n<p>I have a question about a part of the starter kit kernel.  In section 1.b it reads:</p>\n\n<blockquote>\n  <p>We'll give you a few external resources for these events in a companion Kernel, if you are determined to augment the training set. </p>\n</blockquote>\n\n<p>Has this companion kernel been released?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 409018,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-10-23T18:12:26.357000",
          "content": "<p>Hi @CPMP,</p>\n\n<p>It has - it went live with the challenge. I expect it'll get noticed more as people get further into the challenge. The resources are listed in a table at the bottom.</p>\n\n<p><a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-classification-demo\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-classification-demo</a></p>\n\n<p>Best,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 409029,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-23T18:29:59.060000",
          "content": "<p>Thanks a lot, look there is a LOT to explore!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 407986,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2018-10-22T05:25:02.087000",
      "content": "<p>Hello Renée and PLAsTiCC team, thank you for holding such an interesting competition.<br>\nbtw, I have a few questions about <code>detected</code> and <code>flux_err, hostgal_photoz_err</code>in the data set. <br><br>\nI found many samples, which satisfy <code>detected == 1</code> &amp; <code>flux &lt; 0</code> at the same time. <br> \nI think it's strange because <code>data_note.pdf</code> says <br>\n<code>If detected = 1, the object's brightness is significantly different at the 3\\sigma level relative to the reference template. This is given as Boolean flag.</code> <br> <br></p>\n\n<p>In addition, I couldn't understand how <code>flux_err, hostgal_photoz_err</code> are caluclated, which may have something to do with <code>detected</code>. <br><br>\nMy Question: <br>\n1. How is <code>flux_err</code> and <code>hostgal_photoz_err</code> caluclated? <br>\n2. How is <code>detected</code> calculated ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 413861,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2018-11-01T16:57:24.877000",
          "content": "<p>I'm wondering if PLAsTiCC team have had chance to look at the comment above. thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 413868,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-01T17:06:30.680000",
          "content": "<blockquote>\n  <p>detected == 1 &amp; flux &lt; 0 at the same time. </p>\n</blockquote>\n\n<p>This can be true is flux is so low that it is more than 3 std below the reference level.  That's how I understand this anyway, a confirmation would be great.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 413876,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2018-11-01T17:32:03.130000",
          "content": "<p>Hi CPMP, thanks for the reply. <br>\nI first thought same as you, but I found some negative flux samples which is not 3 std below the reference level. \nI also found some positive flux samples which is not 3 std above the reference level. please see the images.\n<br>\n<a href=\"https://imgur.com/srPyhuC\">https://imgur.com/srPyhuC</a> <br>\n<a href=\"https://imgur.com/vW7amPE\">https://imgur.com/vW7amPE</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 414045,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-02T02:16:45.833000",
          "content": "<p>my understanding is that: imagine you have a blinking star with average brightness above 3 sigma from the reference level. But at some moments it totally fades and the flux drops below reference levels, also by 3 sigma or more...  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 414389,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2018-11-02T16:37:17.483000",
          "content": "<p>hmm, so the <code>detected</code> is dependent on the other rows? Anyway, we need to wait for the Host's comment for this problem.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 414462,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-02T20:09:55.043000",
          "content": "<p>yes, I am not an astronomer, so Host comments on this would be much better. It would be also nice to know how the flux is normalized, is the \"reference template\" the same for all data in train and the same as for the test. A bit more details on how it's done</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415280,
          "author_name": "Andry",
          "author_url": "",
          "post_date": "2018-11-04T20:02:22.260000",
          "content": "<p>I guess flux_err were calculated using error images/frames [poisson noise (==sqrt(flux/gain)) + readnoise], and detected == 1 means brightness is &gt; 3 std above the background (source) in the differential image (Img - Img_ref). Detected == 1 doesn't mean flux &gt; flux_err×3, it means the diff flux is at 3 std level of the background noise.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 415907,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-05T22:19:04.517000",
          "content": "<blockquote>\n  <p><strong>mamasinkgs wrote</strong></p>\n  \n  <p>I found many samples, which satisfy <code>detected == 1</code> &amp; <code>flux &amp;lt; 0</code> at the same time. <br> </p>\n</blockquote>\n\n<p>The detection logic is based on <code>abs(SNR)</code>, so negative fluxes can result in a detection.</p>\n\n<blockquote>\n  <ol>\n  <li>How is <code>flux_err</code> and <code>hostgal_photoz_err</code> caluclated? <br></li>\n  </ol>\n</blockquote>\n\n<p>The reported <code>flux_err</code> is a statistical uncertainty from the number of (simulated) photo-electrons read from the CCDs, added in quadrature with the uncertainty arising from correcting the <code>flux</code> for the Milky Way reddening <code>mwebv</code> We aply the Milky Way reddening correction because it's likely the first pre-processing step that astronomers will use, and we felt that this was specialized domain knowledge that contestants without an astronomy background cannot be expected to know. </p>\n\n<p>Source detection algorithms are more complex. The <code>detected</code> flag is set with respect to efficiency curves that are determined by injecting fake sources with known brightness into images, and determining the efficiency of recovering these sources. The flux level at which to set <code>detected = 1</code> is determined from these curves, such that on average sources with S/N &gt;= 3 are treated as real. So <code>detected = 1</code> reflects the the image properties, not just the properties of a single source, and there may be small statistical fluctuations for any given observation. </p>\n\n<p><code>hostgal_photoz_err</code> is much more complicated (there's numerous techniques to determine the photo-z, some of which are ML-based) and we'll discuss this in papers after the end of the challenge. Photometric redshifts might be an interesting subject for a challenge in its own right. </p>\n\n<p>Best,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 416033,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-06T04:07:33.370000",
          "content": "<p>Thanks for the explanations.</p>\n\n<blockquote>\n  <p>a statistical uncertainty </p>\n</blockquote>\n\n<p>Is it the standard deviation of the error, twice the std, something else?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416183,
          "author_name": "mamas",
          "author_url": "",
          "post_date": "2018-11-06T10:50:46.863000",
          "content": "<p>Thanks for the explanations !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416481,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-06T17:47:48.560000",
          "content": "<p>Hi @CPMP,</p>\n\n<blockquote>\n  <p>Is it the standard deviation of the error, twice the std, something else?</p>\n</blockquote>\n\n<p>One standard deviation for <code>flux_err</code></p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 416488,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-06T18:14:05.243000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416597,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-07T00:01:14.567000",
          "content": "<p>Thank you for the update! <code>The flux level at which to set detected = 1 is determined from these curves, such that on average sources with S/N &gt;= 3 are treated as real. So detected = 1 reflects the image properties</code> -- does it mean that when detected = 0 from images, such points are usually disregarded by astronomers ?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416668,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-07T03:54:16.990000",
          "content": "<p>Hi Blonde,</p>\n\n<blockquote>\n  <p>does it mean that when detected = 0 from images, such points are usually disregarded by astronomers ? </p>\n</blockquote>\n\n<p>There's lots of different approaches - some groups disregard these \"non-detections\",  while others do not. </p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417040,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2018-11-07T16:36:21.333000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422286,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-16T02:39:44.777000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 422333,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-16T04:51:18.527000",
          "content": "<p>No it is not log scale, I asked the same question here: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69432#latest-409324\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69432#latest-409324</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422338,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-16T04:56:12.623000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 423002,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-17T08:34:13.067000",
          "content": "<p>Desktop.  I fried my macbook pro with kaggle competitions, I'm no longer using it for long, cpu intensive runs.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 404992,
      "author_name": " Igor Krasovskiy",
      "author_url": "",
      "post_date": "2018-10-16T17:52:34.700000",
      "content": "<p>Thanks for links.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "397138": "Hi Kagglers\n\nWe are so pleased that so many of you are already interested in the challenge, that mimics the kinds of data we are going to get from LSST. We have tried to give you as much information as you need for the challenge, and a bit of explanation about the astronomical concepts, but you don't need to be guided too strongly by any of that!&nbsp;\n\nWe hope you are getting familiar with the data through the starter kit here:&nbsp;&nbsp;https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\nIt contains demos of the data and classification algorithms, and gives a sense for what has been tried before.\n\nWe also have shared the data&nbsp;note&nbsp;with the public and astronomers on the arXiv too: https://arxiv.org/abs/1810.00001\nwhere lots of astronomy papers are shared (it is a great place to learn about astronomy more generally).\n\nThere has been lots of talk about metrics on the discussion boards. FYI we submitted a paper which talks a little bit about metrics for challenges like PLAsTiCC: https://arxiv.org/abs/1809.11145 \n\nThe  release note is geared for people without astronomy training, while the latter is more technical, and geared for astronomers/statistical enthusiasts.\n\nHope you enjoy them both, and the challenge.\n\n- Renée for the PLAsTiCC team",
    "399104": "Hi Renée,\n\nMay I ask you a question:\n\nclass_99 objects - can they actually be objects of different nature, the only thing which is in common for them is that we didn't see them in train set? This way I mean they are not a real class.",
    "399261": "hi Anna, what is means is that because LSST will be a new telescope which is better than ones we have at the moment, we can expect LSST to find objects that we haven't found before.  So we'll have no information on what the objects might be, including their class.",
    "414554": "Hi, Is it allowed to use pre-trained model's weights in this competition ?",
    "416974": "Is it possible to  upload the data from this git repo on Kaggle? There is no license file:\n\nhttps://github.com/lsst/throughputs/tree/master/baseline\n\nI am asking because I'd like to explain how I managed to reproduce the throughput curves using only pandas and matplotlib.  It may help participants understand what is measured.\n\nMy throughput curves:\n\n![throughput curves][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416974/10631/throughput.png",
    "439515": "@kyleboone @gsnarayan\n\nTo clarify: are the given data corrected for the filters opacity? Yes/No ? ",
    "436696": "I was unable to unpack test_set.csv.zip on two different computers, one of them quite new,  with multiple download attempts. The issue appears to be the size of the file, since I can unpack everything else. I there a way to download it in multiple smaller subsets?",
    "418562": "I have a question about flux values. \nAs mentioned above by CPMP, each filters have specific throughput characteristics. Are flux values in the competitions data already corrected from these effect, or raw values from censor?\nIn the latter case, I think we have to correct those to able to equaly treat each passband. ",
    "408520": "Hi, this looks like a quite challenging competition, thanks for hosting it!  \n\nI have a question about a part of the starter kit kernel.  In section 1.b it reads:\n\n&gt; We'll give you a few external resources for these events in a companion Kernel, if you are determined to augment the training set. \n\nHas this companion kernel been released?",
    "407986": "Hello Renée and PLAsTiCC team, thank you for holding such an interesting competition.<br>\nbtw, I have a few questions about `detected` and `flux_err, hostgal_photoz_err`in the data set. <br><br>\nI found many samples, which satisfy `detected == 1` &amp; `flux &lt; 0` at the same time. <br> \nI think it's strange because `data_note.pdf` says <br>\n`If detected = 1, the object's brightness is significantly different at the 3\\sigma level relative to the reference template. This is given as Boolean flag.` <br> <br>\n\nIn addition, I couldn't understand how `flux_err, hostgal_photoz_err` are caluclated, which may have something to do with `detected`. <br><br>\nMy Question: <br>\n1. How is `flux_err` and `hostgal_photoz_err` caluclated? <br>\n2. How is `detected` calculated ?",
    "422286": "",
    "404992": "Thanks for links."
  }
}