{
  "id": 70783,
  "title": "Cesium Features on the Test Set",
  "url": "/competitions/PLAsTiCC-2018/discussion/70783",
  "author_name": "",
  "post_date": "2018-11-07T10:33:17.095702200Z",
  "votes": 31,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Link: <a href=\"https://www.kaggle.com/mithrillion/plasticc-features/\">https://www.kaggle.com/mithrillion/plasticc-features/</a></p>\n\n<p>Probably a total waste of time, but finally managed to get a full table of features for the test set as recommended by <a href=\"https://arxiv.org/abs/1101.1959\">https://arxiv.org/abs/1101.1959</a> .</p>\n\n<p>Please share what features you find useful / redundant if you feel generous!</p>\n\n<p>Personally, I find many features that characterise whether a curve has a peak (like skew) or whether it is periodic (explained variance by frequency) quite useful, but the more detailed information about the curves (like exact frequency) not so much.</p>",
  "messages": [
    {
      "id": "416828",
      "postDate": "11/07/2018 10:33:17",
      "content": "<p>Link: <a href=\"https://www.kaggle.com/mithrillion/plasticc-features/\">https://www.kaggle.com/mithrillion/plasticc-features/</a></p>\n\n<p>Probably a total waste of time, but finally managed to get a full table of features for the test set as recommended by <a href=\"https://arxiv.org/abs/1101.1959\">https://arxiv.org/abs/1101.1959</a> .</p>\n\n<p>Please share what features you find useful / redundant if you feel generous!</p>\n\n<p>Personally, I find many features that characterise whether a curve has a peak (like skew) or whether it is periodic (explained variance by frequency) quite useful, but the more detailed information about the curves (like exact frequency) not so much.</p>",
      "rawMarkdown": "Link: https://www.kaggle.com/mithrillion/plasticc-features/\n\nProbably a total waste of time, but finally managed to get a full table of features for the test set as recommended by https://arxiv.org/abs/1101.1959 .\n\nPlease share what features you find useful / redundant if you feel generous!\n\nPersonally, I find many features that characterise whether a curve has a peak (like skew) or whether it is periodic (explained variance by frequency) quite useful, but the more detailed information about the curves (like exact frequency) not so much.",
      "votes": null
    },
    {
      "id": "416951",
      "postDate": "11/07/2018 13:59:14",
      "content": "<p>Thank you, this is really generous. Would you mind uploading the same features for the training set as well please?</p>",
      "rawMarkdown": "Thank you, this is really generous. Would you mind uploading the same features for the training set as well please?",
      "votes": null
    },
    {
      "id": "417010",
      "postDate": "11/07/2018 15:59:48",
      "content": "<p>Did you compute these passband per passband?</p>",
      "rawMarkdown": "Did you compute these passband per passband?",
      "votes": null
    },
    {
      "id": "417036",
      "postDate": "11/07/2018 16:33:01",
      "content": "<p>Yes he did. Would you recommend otherwise?</p>",
      "rawMarkdown": "Yes he did. Would you recommend otherwise?",
      "votes": null
    },
    {
      "id": "417054",
      "postDate": "11/07/2018 16:57:15",
      "content": "<p>I'm just trying to understand what is in this dataset as it is not described really.  The page he poitns to has only one set of features, not one per passband, hence my question.</p>",
      "rawMarkdown": "I'm just trying to understand what is in this dataset as it is not described really.  The page he poitns to has only one set of features, not one per passband, hence my question.",
      "votes": null
    },
    {
      "id": "417130",
      "postDate": "11/07/2018 20:11:56",
      "content": "<p>Sure. Updating soon.</p>",
      "rawMarkdown": "Sure. Updating soon.",
      "votes": null
    },
    {
      "id": "417132",
      "postDate": "11/07/2018 20:14:03",
      "content": "<p>Yes, because cesium can only do so much... There are more sophisticated algorithms which for instance can calculate a common period using all bands, but it might not be worth it to do the expensive calculation for all examples. I do use them for insights though.</p>",
      "rawMarkdown": "Yes, because cesium can only do so much... There are more sophisticated algorithms which for instance can calculate a common period using all bands, but it might not be worth it to do the expensive calculation for all examples. I do use them for insights though.",
      "votes": null
    },
    {
      "id": "417134",
      "postDate": "11/07/2018 20:14:20",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "417401",
      "postDate": "11/08/2018 07:46:16",
      "content": "<p>Thank you! </p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "417518",
      "postDate": "11/08/2018 12:06:42",
      "content": "<p>did you calculate ALL of them ? How much time did it take for the test? It's a lot to ask but if you can share your part of the code for adding those features calculated for the test it will be helpful. I have problems with going through my bugs (new to python) trying to calculate test features in chunks using Oliver's kernel </p>",
      "rawMarkdown": "did you calculate ALL of them ? How much time did it take for the test? It's a lot to ask but if you can share your part of the code for adding those features calculated for the test it will be helpful. I have problems with going through my bugs (new to python) trying to calculate test features in chunks using Oliver's kernel",
      "votes": null
    },
    {
      "id": "417519",
      "postDate": "11/08/2018 12:07:58",
      "content": "<p>arxiv link is dead, by the way</p>",
      "rawMarkdown": "arxiv link is dead, by the way",
      "votes": null
    },
    {
      "id": "417546",
      "postDate": "11/08/2018 12:42:16",
      "content": "<p>Remove the trailing dot ;)</p>\n\n<p><a href=\"https://arxiv.org/abs/1101.1959\">https://arxiv.org/abs/1101.1959</a></p>",
      "rawMarkdown": "Remove the trailing dot ;)\n\nhttps://arxiv.org/abs/1101.1959",
      "votes": null
    },
    {
      "id": "417658",
      "postDate": "11/08/2018 15:35:24",
      "content": "<p>my God! I am hopeless... what i am doing here :)?</p>",
      "rawMarkdown": "my God! I am hopeless... what i am doing here :)?",
      "votes": null
    },
    {
      "id": "417666",
      "postDate": "11/08/2018 15:53:07",
      "content": "<p>Sorry, just fixed. It sometimes works without an explicit space, sometimes doesn't work. I just can't figure out why...</p>",
      "rawMarkdown": "Sorry, just fixed. It sometimes works without an explicit space, sometimes doesn't work. I just can't figure out why...",
      "votes": null
    },
    {
      "id": "418456",
      "postDate": "11/09/2018 23:18:05",
      "content": "<p>Hey Mithrillion, thanks for uploading! I think I may have found some discrepancies between the <code>freq1_freq</code> feature in the uploaded datasets and the <code>freq1_freq</code> you calculated in your light curve kernel.  More info here: <a href=\"https://www.kaggle.com/mithrillion/plasticc-features/discussion/71077\">https://www.kaggle.com/mithrillion/plasticc-features/discussion/71077</a></p>",
      "rawMarkdown": "Hey Mithrillion, thanks for uploading! I think I may have found some discrepancies between the `freq1_freq` feature in the uploaded datasets and the `freq1_freq` you calculated in your light curve kernel.  More info here: https://www.kaggle.com/mithrillion/plasticc-features/discussion/71077",
      "votes": null
    },
    {
      "id": "418490",
      "postDate": "11/10/2018 01:14:30",
      "content": "<p><strong>EDIT</strong>: I think I just found the reason. Apparently for some cases, supplying observation error estimates to Cesium degrades the quality of frequency estimates. Who would have known...</p>\n\n<p>Thanks for bringing this to my attention. It's very interesting. Apparently many of the results do match up but for some of them the difference is quite large. AFAIK the only difference between the light curve kernel and the uploaded dataset is that in my light curve characteristics kernel, all light curves are individually standardised and no error values are provided to cesium, whereas for the dataset, the raw values are used and errors are supplied. It appears the standardised / no error version works better for some cases. I would assume that standardisation should not matter at all when doing the calculation on individual passbands separately, but I could be wrong. I will try to find the cause of this mismatch.</p>",
      "rawMarkdown": "**EDIT**: I think I just found the reason. Apparently for some cases, supplying observation error estimates to Cesium degrades the quality of frequency estimates. Who would have known...\n\nThanks for bringing this to my attention. It's very interesting. Apparently many of the results do match up but for some of them the difference is quite large. AFAIK the only difference between the light curve kernel and the uploaded dataset is that in my light curve characteristics kernel, all light curves are individually standardised and no error values are provided to cesium, whereas for the dataset, the raw values are used and errors are supplied. It appears the standardised / no error version works better for some cases. I would assume that standardisation should not matter at all when doing the calculation on individual passbands separately, but I could be wrong. I will try to find the cause of this mismatch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 416951,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "11/07/2018 13:59:14",
      "content": "<p>Thank you, this is really generous. Would you mind uploading the same features for the training set as well please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 417130,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/07/2018 20:11:56",
          "content": "<p>Sure. Updating soon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417134,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "11/07/2018 20:14:20",
          "content": "<p>thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417010,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/07/2018 15:59:48",
      "content": "<p>Did you compute these passband per passband?</p>",
      "votes": null,
      "replies": [
        {
          "id": 417036,
          "author_name": "iprapas",
          "author_url": "",
          "post_date": "11/07/2018 16:33:01",
          "content": "<p>Yes he did. Would you recommend otherwise?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417054,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/07/2018 16:57:15",
          "content": "<p>I'm just trying to understand what is in this dataset as it is not described really.  The page he poitns to has only one set of features, not one per passband, hence my question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417132,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/07/2018 20:14:03",
          "content": "<p>Yes, because cesium can only do so much... There are more sophisticated algorithms which for instance can calculate a common period using all bands, but it might not be worth it to do the expensive calculation for all examples. I do use them for insights though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417401,
      "author_name": "andreusancho",
      "author_url": "",
      "post_date": "11/08/2018 07:46:16",
      "content": "<p>Thank you! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417518,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "11/08/2018 12:06:42",
      "content": "<p>did you calculate ALL of them ? How much time did it take for the test? It's a lot to ask but if you can share your part of the code for adding those features calculated for the test it will be helpful. I have problems with going through my bugs (new to python) trying to calculate test features in chunks using Oliver's kernel </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417519,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "11/08/2018 12:07:58",
      "content": "<p>arxiv link is dead, by the way</p>",
      "votes": null,
      "replies": [
        {
          "id": 417546,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/08/2018 12:42:16",
          "content": "<p>Remove the trailing dot ;)</p>\n\n<p><a href=\"https://arxiv.org/abs/1101.1959\">https://arxiv.org/abs/1101.1959</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417658,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/08/2018 15:35:24",
          "content": "<p>my God! I am hopeless... what i am doing here :)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417666,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/08/2018 15:53:07",
          "content": "<p>Sorry, just fixed. It sometimes works without an explicit space, sometimes doesn't work. I just can't figure out why...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418456,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "11/09/2018 23:18:05",
      "content": "<p>Hey Mithrillion, thanks for uploading! I think I may have found some discrepancies between the <code>freq1_freq</code> feature in the uploaded datasets and the <code>freq1_freq</code> you calculated in your light curve kernel.  More info here: <a href=\"https://www.kaggle.com/mithrillion/plasticc-features/discussion/71077\">https://www.kaggle.com/mithrillion/plasticc-features/discussion/71077</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 418490,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "11/10/2018 01:14:30",
          "content": "<p><strong>EDIT</strong>: I think I just found the reason. Apparently for some cases, supplying observation error estimates to Cesium degrades the quality of frequency estimates. Who would have known...</p>\n\n<p>Thanks for bringing this to my attention. It's very interesting. Apparently many of the results do match up but for some of them the difference is quite large. AFAIK the only difference between the light curve kernel and the uploaded dataset is that in my light curve characteristics kernel, all light curves are individually standardised and no error values are provided to cesium, whereas for the dataset, the raw values are used and errors are supplied. It appears the standardised / no error version works better for some cases. I would assume that standardisation should not matter at all when doing the calculation on individual passbands separately, but I could be wrong. I will try to find the cause of this mismatch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "416828": "Link: https://www.kaggle.com/mithrillion/plasticc-features/\n\nProbably a total waste of time, but finally managed to get a full table of features for the test set as recommended by https://arxiv.org/abs/1101.1959 .\n\nPlease share what features you find useful / redundant if you feel generous!\n\nPersonally, I find many features that characterise whether a curve has a peak (like skew) or whether it is periodic (explained variance by frequency) quite useful, but the more detailed information about the curves (like exact frequency) not so much.",
    "416951": "Thank you, this is really generous. Would you mind uploading the same features for the training set as well please?",
    "417010": "Did you compute these passband per passband?",
    "417036": "Yes he did. Would you recommend otherwise?",
    "417054": "I'm just trying to understand what is in this dataset as it is not described really.  The page he poitns to has only one set of features, not one per passband, hence my question.",
    "417130": "Sure. Updating soon.",
    "417132": "Yes, because cesium can only do so much... There are more sophisticated algorithms which for instance can calculate a common period using all bands, but it might not be worth it to do the expensive calculation for all examples. I do use them for insights though.",
    "417134": "thank you!",
    "417401": "Thank you!",
    "417518": "did you calculate ALL of them ? How much time did it take for the test? It's a lot to ask but if you can share your part of the code for adding those features calculated for the test it will be helpful. I have problems with going through my bugs (new to python) trying to calculate test features in chunks using Oliver's kernel",
    "417519": "arxiv link is dead, by the way",
    "417546": "Remove the trailing dot ;)\n\nhttps://arxiv.org/abs/1101.1959",
    "417658": "my God! I am hopeless... what i am doing here :)?",
    "417666": "Sorry, just fixed. It sometimes works without an explicit space, sometimes doesn't work. I just can't figure out why...",
    "418456": "Hey Mithrillion, thanks for uploading! I think I may have found some discrepancies between the `freq1_freq` feature in the uploaded datasets and the `freq1_freq` you calculated in your light curve kernel.  More info here: https://www.kaggle.com/mithrillion/plasticc-features/discussion/71077",
    "418490": "**EDIT**: I think I just found the reason. Apparently for some cases, supplying observation error estimates to Cesium degrades the quality of frequency estimates. Who would have known...\n\nThanks for bringing this to my attention. It's very interesting. Apparently many of the results do match up but for some of them the difference is quite large. AFAIK the only difference between the light curve kernel and the uploaded dataset is that in my light curve characteristics kernel, all light curves are individually standardised and no error values are provided to cesium, whereas for the dataset, the raw values are used and errors are supplied. It appears the standardised / no error version works better for some cases. I would assume that standardisation should not matter at all when doing the calculation on individual passbands separately, but I could be wrong. I will try to find the cause of this mismatch."
  },
  "source": "meta"
}