{
  "id": 71708,
  "title": "Simulated data",
  "url": "/competitions/PLAsTiCC-2018/discussion/71708",
  "author_name": "Michal Haltuf",
  "post_date": "2018-11-15T19:54:56.373000",
  "votes": 9,
  "comment_count": 11,
  "views": 0,
  "content": "<p>So I was reading the official data_note.pdf for the first time today and was quite surprised by the information, that <strong>all the dataset is simulated</strong> - i.e. artificially created, computer generated.\nDo I understand it correctly?</p>\n\n<p>Is there any way we could leverage this knowledge? (e.g. for data augmentation)</p>\n\n<p>How does one create simulated data of astronomical time-series? I noticed it's quite a common practice in training supernovae classifiers, but haven't found any paper that would be detailed and specific enough, how to create your own astronomical time-serie simulator for different object classes.</p>",
  "messages": [
    {
      "id": 422111,
      "postDate": "2018-11-15T19:54:56.373Z",
      "content": "<p>So I was reading the official data_note.pdf for the first time today and was quite surprised by the information, that <strong>all the dataset is simulated</strong> - i.e. artificially created, computer generated.\nDo I understand it correctly?</p>\n\n<p>Is there any way we could leverage this knowledge? (e.g. for data augmentation)</p>\n\n<p>How does one create simulated data of astronomical time-series? I noticed it's quite a common practice in training supernovae classifiers, but haven't found any paper that would be detailed and specific enough, how to create your own astronomical time-serie simulator for different object classes.</p>",
      "rawMarkdown": "So I was reading the official data_note.pdf for the first time today and was quite surprised by the information, that **all the dataset is simulated** - i.e. artificially created, computer generated.\nDo I understand it correctly?\n\nIs there any way we could leverage this knowledge? (e.g. for data augmentation)\n\nHow does one create simulated data of astronomical time-series? I noticed it's quite a common practice in training supernovae classifiers, but haven't found any paper that would be detailed and specific enough, how to create your own astronomical time-serie simulator for different object classes.",
      "votes": 9
    },
    {
      "id": 422219,
      "postDate": "2018-11-15T23:37:22.983Z",
      "content": "<p>I think the training dataset is small to simulate the real-life challenge: having a limited amount of real data to classify a vast amount of new data they are expecting in 2020</p>",
      "rawMarkdown": "I think the training dataset is small to simulate the real-life challenge: having a limited amount of real data to classify a vast amount of new data they are expecting in 2020",
      "votes": 3,
      "replies": [
        {
          "id": 422239,
          "postDate": "2018-11-16T00:12:25.010Z",
          "content": "<p>Correct.</p>",
          "rawMarkdown": "Correct."
        }
      ]
    },
    {
      "id": 422213,
      "postDate": "2018-11-15T23:27:15.137Z",
      "content": "<p>If the answer exists anywhere, it would be in this repo collection: <a href=\"https://github.com/lsst\">https://github.com/lsst</a></p>\n\n<p>The competition organizers have made note that they would be releasing a paper post-competition outlining the data construction strategy though, so it'd be a big misstep for them to have jumped the gun and already released source code for that publicly given this is a paid contest.</p>",
      "rawMarkdown": "If the answer exists anywhere, it would be in this repo collection: https://github.com/lsst\n\nThe competition organizers have made note that they would be releasing a paper post-competition outlining the data construction strategy though, so it'd be a big misstep for them to have jumped the gun and already released source code for that publicly given this is a paid contest.",
      "votes": 1
    },
    {
      "id": 422293,
      "postDate": "2018-11-16T02:58:43.003Z",
      "content": "<p>LSST has an open source Light Curve simulator repo on GitHub (<a href=\"https://github.com/lsst/sims_catUtils\">https://github.com/lsst/sims_catUtils</a> ) that was last updated a month ago.  It even includes IPython notebooks with examples.</p>",
      "rawMarkdown": "LSST has an open source Light Curve simulator repo on GitHub (https://github.com/lsst/sims_catUtils ) that was last updated a month ago.  It even includes IPython notebooks with examples.",
      "replies": [
        {
          "id": 422368,
          "postDate": "2018-11-16T06:03:32.260Z",
          "content": "<p>I found this a month ago and decided to:</p>\n\n<ol>\n<li>Not look at it</li>\n<li>Not share the link here to not tease people</li>\n</ol>\n\n<p>Using this would defeat the purpose of the competition IMHO.</p>",
          "rawMarkdown": "I found this a month ago and decided to:\n\n 1. Not look at it\n 2. Not share the link here to not tease people\n\nUsing this would defeat the purpose of the competition IMHO.",
          "votes": 9
        },
        {
          "id": 422465,
          "postDate": "2018-11-16T08:58:14.943Z",
          "content": "<p>agree+1</p>",
          "rawMarkdown": "agree+1",
          "votes": 1
        },
        {
          "id": 422731,
          "postDate": "2018-11-16T17:34:01.740Z",
          "content": "<p>I suspect anyone trying to use that tool would find that generating equivalent data is non-trivial, even for a trained astronomer. For a tiny taste of the complexity involved, remember that to generate realistic light curves you need to specify parameters like the shape of the universe, amount of dark energy in existence, etc.</p>\n\n<p>Based on the level of expertise it took to generate the main competition dataset, I highly doubt that trying to augment the dataset with fresh light curves would be productive. More likely, the tool would spit out plausible seeming but useless curves.</p>",
          "rawMarkdown": "I suspect anyone trying to use that tool would find that generating equivalent data is non-trivial, even for a trained astronomer. For a tiny taste of the complexity involved, remember that to generate realistic light curves you need to specify parameters like the shape of the universe, amount of dark energy in existence, etc.\n\nBased on the level of expertise it took to generate the main competition dataset, I highly doubt that trying to augment the dataset with fresh light curves would be productive. More likely, the tool would spit out plausible seeming but useless curves."
        },
        {
          "id": 422744,
          "postDate": "2018-11-16T17:54:15.450Z",
          "content": "<p>We'd echo Sohier's point, and suggest that trying to augment the data this way is much harder than one might appreciate. Building generative models from the existing training set might prove more fruitful and is definitely of scientific interest to us. </p>\n\n<p>That said, you should feel free to fiddle with this code if you believe it will help you.</p>\n\n<p>[insert evil cackling here]</p>\n\n<p>Cheers,\n-Gautham for the PLAsTiCC team</p>",
          "rawMarkdown": "We'd echo Sohier's point, and suggest that trying to augment the data this way is much harder than one might appreciate. Building generative models from the existing training set might prove more fruitful and is definitely of scientific interest to us. \n\nThat said, you should feel free to fiddle with this code if you believe it will help you.\n\n[insert evil cackling here]\n\nCheers,\n-Gautham for the PLAsTiCC team\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 422166,
      "postDate": "2018-11-15T21:32:27.150Z",
      "content": "<p>I wonder if the dataset is simulated, why is the training dataset so small? (In the beginning I thought this is real data)</p>",
      "rawMarkdown": "I wonder if the dataset is simulated, why is the training dataset so small? (In the beginning I thought this is real data)"
    },
    {
      "id": 422134,
      "postDate": "2018-11-15T20:27:33.250Z",
      "content": "<p>Late, as usual, I've noticed, there are some hints or maybe even answers for my questions, here: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71544\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71544</a> and here: <a href=\"https://www.icts.res.in/program/TASSGW2017/talks\">https://www.icts.res.in/program/TASSGW2017/talks</a></p>",
      "rawMarkdown": "Late, as usual, I've noticed, there are some hints or maybe even answers for my questions, here: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71544 and here: https://www.icts.res.in/program/TASSGW2017/talks"
    },
    {
      "id": 422329,
      "postDate": "2018-11-16T04:30:15.890Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 422219,
      "author_name": "Blonde",
      "author_url": "",
      "post_date": "2018-11-15T23:37:22.983000",
      "content": "<p>I think the training dataset is small to simulate the real-life challenge: having a limited amount of real data to classify a vast amount of new data they are expecting in 2020</p>",
      "votes": 3,
      "replies": [
        {
          "id": 422239,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-11-16T00:12:25.010000",
          "content": "<p>Correct.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422213,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2018-11-15T23:27:15.137000",
      "content": "<p>If the answer exists anywhere, it would be in this repo collection: <a href=\"https://github.com/lsst\">https://github.com/lsst</a></p>\n\n<p>The competition organizers have made note that they would be releasing a paper post-competition outlining the data construction strategy though, so it'd be a big misstep for them to have jumped the gun and already released source code for that publicly given this is a paid contest.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 422293,
      "author_name": "Mike Holcomb",
      "author_url": "",
      "post_date": "2018-11-16T02:58:43.003000",
      "content": "<p>LSST has an open source Light Curve simulator repo on GitHub (<a href=\"https://github.com/lsst/sims_catUtils\">https://github.com/lsst/sims_catUtils</a> ) that was last updated a month ago.  It even includes IPython notebooks with examples.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 422368,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-16T06:03:32.260000",
          "content": "<p>I found this a month ago and decided to:</p>\n\n<ol>\n<li>Not look at it</li>\n<li>Not share the link here to not tease people</li>\n</ol>\n\n<p>Using this would defeat the purpose of the competition IMHO.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 422465,
          "author_name": "khyeh",
          "author_url": "",
          "post_date": "2018-11-16T08:58:14.943000",
          "content": "<p>agree+1</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422731,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2018-11-16T17:34:01.740000",
          "content": "<p>I suspect anyone trying to use that tool would find that generating equivalent data is non-trivial, even for a trained astronomer. For a tiny taste of the complexity involved, remember that to generate realistic light curves you need to specify parameters like the shape of the universe, amount of dark energy in existence, etc.</p>\n\n<p>Based on the level of expertise it took to generate the main competition dataset, I highly doubt that trying to augment the dataset with fresh light curves would be productive. More likely, the tool would spit out plausible seeming but useless curves.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422744,
          "author_name": "Gautham Narayan",
          "author_url": "",
          "post_date": "2018-11-16T17:54:15.450000",
          "content": "<p>We'd echo Sohier's point, and suggest that trying to augment the data this way is much harder than one might appreciate. Building generative models from the existing training set might prove more fruitful and is definitely of scientific interest to us. </p>\n\n<p>That said, you should feel free to fiddle with this code if you believe it will help you.</p>\n\n<p>[insert evil cackling here]</p>\n\n<p>Cheers,\n-Gautham for the PLAsTiCC team</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 422166,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-11-15T21:32:27.150000",
      "content": "<p>I wonder if the dataset is simulated, why is the training dataset so small? (In the beginning I thought this is real data)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422134,
      "author_name": "Michal Haltuf",
      "author_url": "",
      "post_date": "2018-11-15T20:27:33.250000",
      "content": "<p>Late, as usual, I've noticed, there are some hints or maybe even answers for my questions, here: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71544\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71544</a> and here: <a href=\"https://www.icts.res.in/program/TASSGW2017/talks\">https://www.icts.res.in/program/TASSGW2017/talks</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 422329,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-16T04:30:15.890000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422111": "So I was reading the official data_note.pdf for the first time today and was quite surprised by the information, that **all the dataset is simulated** - i.e. artificially created, computer generated.\nDo I understand it correctly?\n\nIs there any way we could leverage this knowledge? (e.g. for data augmentation)\n\nHow does one create simulated data of astronomical time-series? I noticed it's quite a common practice in training supernovae classifiers, but haven't found any paper that would be detailed and specific enough, how to create your own astronomical time-serie simulator for different object classes.",
    "422219": "I think the training dataset is small to simulate the real-life challenge: having a limited amount of real data to classify a vast amount of new data they are expecting in 2020",
    "422213": "If the answer exists anywhere, it would be in this repo collection: https://github.com/lsst\n\nThe competition organizers have made note that they would be releasing a paper post-competition outlining the data construction strategy though, so it'd be a big misstep for them to have jumped the gun and already released source code for that publicly given this is a paid contest.",
    "422293": "LSST has an open source Light Curve simulator repo on GitHub (https://github.com/lsst/sims_catUtils ) that was last updated a month ago.  It even includes IPython notebooks with examples.",
    "422166": "I wonder if the dataset is simulated, why is the training dataset so small? (In the beginning I thought this is real data)",
    "422134": "Late, as usual, I've noticed, there are some hints or maybe even answers for my questions, here: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71544 and here: https://www.icts.res.in/program/TASSGW2017/talks",
    "422329": ""
  }
}