{
  "id": 70712,
  "title": "Source code for a complete solution",
  "url": "/competitions/PLAsTiCC-2018/discussion/70712",
  "author_name": "",
  "post_date": "2018-11-06T17:14:47.122352800Z",
  "votes": 32,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Dear all,</p>\n\n<p>I spent quite a lot of time on this competition, and would like to share my code now. I placed my scripts here:</p>\n\n<p><a href=\"https://github.com/JohannesBuchner/LSST-PLAsTiCC-classification-solution\">https://github.com/JohannesBuchner/LSST-PLAsTiCC-classification-solution</a></p>\n\n<p>I documented there:</p>\n\n<ol>\n<li>Data visualisation</li>\n<li>Data Transform: Making time series features </li>\n<li>Data Transform of test data</li>\n<li>Training models</li>\n<li>Training a hyper-classifier</li>\n<li>Novelty detection</li>\n<li>Blending novelty detection and classifiers</li>\n<li>Submitting to Kaggle</li>\n</ol>\n\n<p>There are some neat features implemented you might find interesting (or amusing), including:</p>\n\n<ul>\n<li>counting dips, peaks and consecutive runs</li>\n<li>log-linear regression and boot-strapped rest-frame fitting, redshift resampling</li>\n<li>Black-body temperature fitting</li>\n</ul>\n\n<p>For the machine learning itself, I tried to keep it simple (CPU-only training, doesn't need a data center).</p>\n\n<p>For some reason I could never go below the 1.79-2.00 range. If you have any comments, please send them my way (here or github issues).</p>\n\n<p>I hope this is useful for some late-comers to get started. All scripts are automated and can use multiple CPUs.</p>\n\n<p>Cheers,\n                 Johannes</p>",
  "messages": [
    {
      "id": "416463",
      "postDate": "11/06/2018 17:14:47",
      "content": "<p>Dear all,</p>\n\n<p>I spent quite a lot of time on this competition, and would like to share my code now. I placed my scripts here:</p>\n\n<p><a href=\"https://github.com/JohannesBuchner/LSST-PLAsTiCC-classification-solution\">https://github.com/JohannesBuchner/LSST-PLAsTiCC-classification-solution</a></p>\n\n<p>I documented there:</p>\n\n<ol>\n<li>Data visualisation</li>\n<li>Data Transform: Making time series features </li>\n<li>Data Transform of test data</li>\n<li>Training models</li>\n<li>Training a hyper-classifier</li>\n<li>Novelty detection</li>\n<li>Blending novelty detection and classifiers</li>\n<li>Submitting to Kaggle</li>\n</ol>\n\n<p>There are some neat features implemented you might find interesting (or amusing), including:</p>\n\n<ul>\n<li>counting dips, peaks and consecutive runs</li>\n<li>log-linear regression and boot-strapped rest-frame fitting, redshift resampling</li>\n<li>Black-body temperature fitting</li>\n</ul>\n\n<p>For the machine learning itself, I tried to keep it simple (CPU-only training, doesn't need a data center).</p>\n\n<p>For some reason I could never go below the 1.79-2.00 range. If you have any comments, please send them my way (here or github issues).</p>\n\n<p>I hope this is useful for some late-comers to get started. All scripts are automated and can use multiple CPUs.</p>\n\n<p>Cheers,\n                 Johannes</p>",
      "rawMarkdown": "Dear all,\n\nI spent quite a lot of time on this competition, and would like to share my code now. I placed my scripts here:\n\nhttps://github.com/JohannesBuchner/LSST-PLAsTiCC-classification-solution\n\nI documented there:\n\n 0. Data visualisation\n 1. Data Transform: Making time series features \n 2. Data Transform of test data\n 3. Training models\n 4. Training a hyper-classifier\n 5. Novelty detection\n 6. Blending novelty detection and classifiers\n 7. Submitting to Kaggle\n\nThere are some neat features implemented you might find interesting (or amusing), including:\n\n * counting dips, peaks and consecutive runs\n * log-linear regression and boot-strapped rest-frame fitting, redshift resampling\n * Black-body temperature fitting\n\nFor the machine learning itself, I tried to keep it simple (CPU-only training, doesn't need a data center).\n\nFor some reason I could never go below the 1.79-2.00 range. If you have any comments, please send them my way (here or github issues).\n\nI hope this is useful for some late-comers to get started. All scripts are automated and can use multiple CPUs.\n\nCheers,\n                 Johannes",
      "votes": null
    },
    {
      "id": "416496",
      "postDate": "11/06/2018 18:36:26",
      "content": "<p>There seems to be a lot of things.  I may be asking for a lot, but can you indicate which file implements which piece of what you describe int he README?</p>",
      "rawMarkdown": "There seems to be a lot of things.  I may be asking for a lot, but can you indicate which file implements which piece of what you describe int he README?",
      "votes": null
    },
    {
      "id": "416500",
      "postDate": "11/06/2018 18:46:58",
      "content": "<p>I think it is written there, just highlight .py.</p>",
      "rawMarkdown": "I think it is written there, just highlight .py.",
      "votes": null
    },
    {
      "id": "416634",
      "postDate": "11/07/2018 02:52:29",
      "content": "<p>This is really nice! However, when you split the data and run \n<code>\nmake -j4 -k chunks/test_set_chunk{0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32}_{std,colorslope,SEDprob}_features.txt\n</code> \nIt throws an error because it cannot find the metafiles for the chunks. </p>",
      "rawMarkdown": "This is really nice! However, when you split the data and run \n```\nmake -j4 -k chunks/test_set_chunk{0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32}_{std,colorslope,SEDprob}_features.txt\n``` \nIt throws an error because it cannot find the metafiles for the chunks.",
      "votes": null
    },
    {
      "id": "416645",
      "postDate": "11/07/2018 03:11:33",
      "content": "<p>Thanks for pointing this out. I fixed the split.sh script. Please try again and let me know.</p>",
      "rawMarkdown": "Thanks for pointing this out. I fixed the split.sh script. Please try again and let me know.",
      "votes": null
    },
    {
      "id": "416799",
      "postDate": "11/07/2018 09:55:47",
      "content": "<p>In your file  <code>deextinct.py</code> I see :</p>\n\n<pre><code>for filtername in 'ugrizy':\n    wave_nm, filter_throughput = numpy.loadtxt('throughputs/baseline/filter_%s.dat' % filtername).transpose()\n</code></pre>\n\n<p>Where are these thruput files?</p>\n\n<p>I also see:</p>\n\n<pre><code>import extinction\n</code></pre>\n\n<p>I don't find it in your git repo.</p>",
      "rawMarkdown": "In your file  `deextinct.py` I see :\n\n    for filtername in 'ugrizy':\n    \twave_nm, filter_throughput = numpy.loadtxt('throughputs/baseline/filter_%s.dat' % filtername).transpose()\n\nWhere are these thruput files?\n\nI also see:\n\n    import extinction\n\nI don't find it in your git repo.",
      "votes": null
    },
    {
      "id": "416864",
      "postDate": "11/07/2018 11:26:57",
      "content": "<p>deextinct.py is not used, you can ignore it. But if you are interested:</p>\n\n<ul>\n<li>extinction is a pypi package <a href=\"https://pypi.org/project/extinction/\">https://pypi.org/project/extinction/</a> </li>\n<li>LSST information <a href=\"https://github.com/lsst/throughputs\">https://github.com/lsst/throughputs</a></li>\n</ul>",
      "rawMarkdown": "deextinct.py is not used, you can ignore it. But if you are interested:\n\n * extinction is a pypi package https://pypi.org/project/extinction/ \n * LSST information https://github.com/lsst/throughputs",
      "votes": null
    },
    {
      "id": "417978",
      "postDate": "11/09/2018 04:22:29",
      "content": "<p>I was able to split the files and the metafiles are now created. However, I still get an error when I run the above command.\nIf I can fix it, I will let you know!\n<code>\nTraceback (most recent call last):\n  File \"make_SED_features.py\", line 32, in &lt;module&gt;\n    e = a.join(b)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/frame.py\", line 6336, in join\n    rsuffix=rsuffix, sort=sort)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/frame.py\", line 6351, in _join_compat\n    suffixes=(lsuffix, rsuffix), sort=sort)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/reshape/merge.py\", line 62, in merge\n    return op.get_result()\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/reshape/merge.py\", line 574, in get_result\n    rdata.items, rsuf)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/internals.py\", line 5244, in items_overlap_with_suffix\n    '{rename}'.format(rename=to_rename))\nValueError: columns overlap but no suffix specified: Index([u'mjd', u'passband', u'flux', u'flux_err', u'detected'], dtype='object')\nMakefile:14: recipe for target 'chunks/test_set_chunk1_SEDprob_features.txt' failed\nmake: *** [chunks/test_set_chunk1_SEDprob_features.txt] Error 1\n</code></p>",
      "rawMarkdown": "I was able to split the files and the metafiles are now created. However, I still get an error when I run the above command.\nIf I can fix it, I will let you know!\n```\nTraceback (most recent call last):\n  File \"make_SED_features.py\", line 32, in",
      "votes": null
    },
    {
      "id": "418047",
      "postDate": "11/09/2018 07:23:35",
      "content": "<p>Is that you: <a href=\"http://astro.puc.cl/~jbuchner/about/\">http://astro.puc.cl/~jbuchner/about/</a> ?</p>",
      "rawMarkdown": "Is that you: http://astro.puc.cl/~jbuchner/about/ ?",
      "votes": null
    },
    {
      "id": "427266",
      "postDate": "11/25/2018 02:22:30",
      "content": "<p>Thank you so much for wonderful work. Can't thank you enough. </p>",
      "rawMarkdown": "Thank you so much for wonderful work. Can't thank you enough.",
      "votes": null
    },
    {
      "id": "435040",
      "postDate": "12/07/2018 11:15:23",
      "content": "<p>Thank you for great work. </p>",
      "rawMarkdown": "Thank you for great work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 416496,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/06/2018 18:36:26",
      "content": "<p>There seems to be a lot of things.  I may be asking for a lot, but can you indicate which file implements which piece of what you describe int he README?</p>",
      "votes": null,
      "replies": [
        {
          "id": 416500,
          "author_name": "jbuchner",
          "author_url": "",
          "post_date": "11/06/2018 18:46:58",
          "content": "<p>I think it is written there, just highlight .py.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 416634,
      "author_name": "jadeflower",
      "author_url": "",
      "post_date": "11/07/2018 02:52:29",
      "content": "<p>This is really nice! However, when you split the data and run \n<code>\nmake -j4 -k chunks/test_set_chunk{0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32}_{std,colorslope,SEDprob}_features.txt\n</code> \nIt throws an error because it cannot find the metafiles for the chunks. </p>",
      "votes": null,
      "replies": [
        {
          "id": 416645,
          "author_name": "jbuchner",
          "author_url": "",
          "post_date": "11/07/2018 03:11:33",
          "content": "<p>Thanks for pointing this out. I fixed the split.sh script. Please try again and let me know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417978,
          "author_name": "jadeflower",
          "author_url": "",
          "post_date": "11/09/2018 04:22:29",
          "content": "<p>I was able to split the files and the metafiles are now created. However, I still get an error when I run the above command.\nIf I can fix it, I will let you know!\n<code>\nTraceback (most recent call last):\n  File \"make_SED_features.py\", line 32, in &lt;module&gt;\n    e = a.join(b)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/frame.py\", line 6336, in join\n    rsuffix=rsuffix, sort=sort)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/frame.py\", line 6351, in _join_compat\n    suffixes=(lsuffix, rsuffix), sort=sort)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/reshape/merge.py\", line 62, in merge\n    return op.get_result()\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/reshape/merge.py\", line 574, in get_result\n    rdata.items, rsuf)\n  File \"/home/sabahat/.local/lib/python2.7/site-packages/pandas/core/internals.py\", line 5244, in items_overlap_with_suffix\n    '{rename}'.format(rename=to_rename))\nValueError: columns overlap but no suffix specified: Index([u'mjd', u'passband', u'flux', u'flux_err', u'detected'], dtype='object')\nMakefile:14: recipe for target 'chunks/test_set_chunk1_SEDprob_features.txt' failed\nmake: *** [chunks/test_set_chunk1_SEDprob_features.txt] Error 1\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 416799,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/07/2018 09:55:47",
      "content": "<p>In your file  <code>deextinct.py</code> I see :</p>\n\n<pre><code>for filtername in 'ugrizy':\n    wave_nm, filter_throughput = numpy.loadtxt('throughputs/baseline/filter_%s.dat' % filtername).transpose()\n</code></pre>\n\n<p>Where are these thruput files?</p>\n\n<p>I also see:</p>\n\n<pre><code>import extinction\n</code></pre>\n\n<p>I don't find it in your git repo.</p>",
      "votes": null,
      "replies": [
        {
          "id": 416864,
          "author_name": "jbuchner",
          "author_url": "",
          "post_date": "11/07/2018 11:26:57",
          "content": "<p>deextinct.py is not used, you can ignore it. But if you are interested:</p>\n\n<ul>\n<li>extinction is a pypi package <a href=\"https://pypi.org/project/extinction/\">https://pypi.org/project/extinction/</a> </li>\n<li>LSST information <a href=\"https://github.com/lsst/throughputs\">https://github.com/lsst/throughputs</a></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418047,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/09/2018 07:23:35",
      "content": "<p>Is that you: <a href=\"http://astro.puc.cl/~jbuchner/about/\">http://astro.puc.cl/~jbuchner/about/</a> ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 427266,
      "author_name": "liger1776",
      "author_url": "",
      "post_date": "11/25/2018 02:22:30",
      "content": "<p>Thank you so much for wonderful work. Can't thank you enough. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 435040,
      "author_name": "tajlos",
      "author_url": "",
      "post_date": "12/07/2018 11:15:23",
      "content": "<p>Thank you for great work. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "416463": "Dear all,\n\nI spent quite a lot of time on this competition, and would like to share my code now. I placed my scripts here:\n\nhttps://github.com/JohannesBuchner/LSST-PLAsTiCC-classification-solution\n\nI documented there:\n\n 0. Data visualisation\n 1. Data Transform: Making time series features \n 2. Data Transform of test data\n 3. Training models\n 4. Training a hyper-classifier\n 5. Novelty detection\n 6. Blending novelty detection and classifiers\n 7. Submitting to Kaggle\n\nThere are some neat features implemented you might find interesting (or amusing), including:\n\n * counting dips, peaks and consecutive runs\n * log-linear regression and boot-strapped rest-frame fitting, redshift resampling\n * Black-body temperature fitting\n\nFor the machine learning itself, I tried to keep it simple (CPU-only training, doesn't need a data center).\n\nFor some reason I could never go below the 1.79-2.00 range. If you have any comments, please send them my way (here or github issues).\n\nI hope this is useful for some late-comers to get started. All scripts are automated and can use multiple CPUs.\n\nCheers,\n                 Johannes",
    "416496": "There seems to be a lot of things.  I may be asking for a lot, but can you indicate which file implements which piece of what you describe int he README?",
    "416500": "I think it is written there, just highlight .py.",
    "416634": "This is really nice! However, when you split the data and run \n```\nmake -j4 -k chunks/test_set_chunk{0,1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32}_{std,colorslope,SEDprob}_features.txt\n``` \nIt throws an error because it cannot find the metafiles for the chunks.",
    "416645": "Thanks for pointing this out. I fixed the split.sh script. Please try again and let me know.",
    "416799": "In your file  `deextinct.py` I see :\n\n    for filtername in 'ugrizy':\n    \twave_nm, filter_throughput = numpy.loadtxt('throughputs/baseline/filter_%s.dat' % filtername).transpose()\n\nWhere are these thruput files?\n\nI also see:\n\n    import extinction\n\nI don't find it in your git repo.",
    "416864": "deextinct.py is not used, you can ignore it. But if you are interested:\n\n * extinction is a pypi package https://pypi.org/project/extinction/ \n * LSST information https://github.com/lsst/throughputs",
    "417978": "I was able to split the files and the metafiles are now created. However, I still get an error when I run the above command.\nIf I can fix it, I will let you know!\n```\nTraceback (most recent call last):\n  File \"make_SED_features.py\", line 32, in",
    "418047": "Is that you: http://astro.puc.cl/~jbuchner/about/ ?",
    "427266": "Thank you so much for wonderful work. Can't thank you enough.",
    "435040": "Thank you for great work."
  },
  "source": "meta"
}