{
  "id": 69696,
  "title": "Use causality",
  "url": "/competitions/PLAsTiCC-2018/discussion/69696",
  "author_name": "CPMP",
  "post_date": "2018-10-26T05:33:23.734000",
  "votes": 23,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I see that some kernels use features based solely on when flux are measured, for instance <code>df['mjd_diff'] = df['mjd_max'] - df['mjd_min']</code></p>\n\n<p>Think of it.  How could the timing of measure influence the nature of the source we measure?  The flux we measure was emitted before it is measured.  There is no way the measure can influence the source.  </p>\n\n<p>Given there cannot be a causality effect, I am not using theses features.</p>\n\n<p>Edited.  Added 'solely' to make the point clearer.</p>",
  "messages": [
    {
      "id": 410538,
      "postDate": "2018-10-26T08:21:34.327Z",
      "content": "<blockquote>\n  <p>I see that some kernels use features based on when flux are measured, ... I am not using theses features.</p>\n</blockquote>\n\n<p>@CPMP, I bet you will ;)</p>\n\n<pre><code>dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n</code></pre>\n\n<p>is a great feature to separate \"one event\" objects as supernovae from \"cyclic event\" objects as cepheids.</p>",
      "rawMarkdown": "&gt; I see that some kernels use features based on when flux are measured, ... I am not using theses features.\n\n@CPMP, I bet you will ;)\n\n    dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n\nis a great feature to separate \"one event\" objects as supernovae from \"cyclic event\" objects as cepheids.",
      "votes": 41,
      "replies": [
        {
          "id": 410573,
          "postDate": "2018-10-26T09:33:17.593Z",
          "content": "<p>This one depends on what you measure, I can use it.  Thanks for sharing.  I edited my post to make it clearer.</p>",
          "rawMarkdown": "This one depends on what you measure, I can use it.  Thanks for sharing.  I edited my post to make it clearer.",
          "votes": 1
        },
        {
          "id": 411119,
          "postDate": "2018-10-27T13:04:17.527Z",
          "content": "<p>Grzegorz, your feature gave me a significant boost of about 0.07,  half of what I gained today,.  Thanks a lot for having shared it. I am sure others got a boost with it too!</p>",
          "rawMarkdown": "Grzegorz, your feature gave me a significant boost of about 0.07,  half of what I gained today,.  Thanks a lot for having shared it. I am sure others got a boost with it too!",
          "votes": 3
        },
        {
          "id": 411155,
          "postDate": "2018-10-27T14:18:16.533Z",
          "content": "<p>That's what I've found in local CV. Hope it will translate to LB ;-)</p>\n\n<p>A big thank you for sharing with us <a href=\"/sionek\">@sionek</a></p>\n\n<p>UPDATE: I got a 0.1 LB improvement</p>",
          "rawMarkdown": "That's what I've found in local CV. Hope it will translate to LB ;-)\n\nA big thank you for sharing with us @sionek\n\nUPDATE: I got a 0.1 LB improvement",
          "votes": 4
        },
        {
          "id": 415791,
          "postDate": "2018-11-05T17:45:11.250Z",
          "content": "<p>Thank you, this feature really helps a lot to my model!</p>",
          "rawMarkdown": "Thank you, this feature really helps a lot to my model!"
        },
        {
          "id": 419065,
          "postDate": "2018-11-11T07:02:24.720Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 419072,
          "postDate": "2018-11-11T07:29:02.383Z",
          "content": "<p>That's using python table I think.  you can do something similar with pandas.groupby().  Something like:</p>\n\n<pre><code>train['mjd_min'] = train.groupby('object_id').mjd.transform('min')\ntrain['mjd_max'] = train.groupby('object_id').mjd.transform('max')\ntrain['mjd_diff']  = train['mjd_max']  - train['mjd_min'] \ndel train['mjd_min'] , train['mjd_max']\n</code></pre>",
          "rawMarkdown": "That's using python table I think.  you can do something similar with pandas.groupby().  Something like:\n\n    train['mjd_min'] = train.groupby('object_id').mjd.transform('min')\n    train['mjd_max'] = train.groupby('object_id').mjd.transform('max')\n    train['mjd_diff']  = train['mjd_max']  - train['mjd_min'] \n    del train['mjd_min'] , train['mjd_max']",
          "votes": 2
        },
        {
          "id": 419989,
          "postDate": "2018-11-12T22:09:12.880Z",
          "content": "<p>Just for everyone's general knowledge:</p>\n\n<p><code>train['mjd_diff']  = train.groupby('object_id').mjd.apply(lambda x: x.ptp())</code></p>\n\n<p>where <code>ptp</code> stands for \"peak to peak\".</p>",
          "rawMarkdown": "Just for everyone's general knowledge:\n\n`train['mjd_diff']  = train.groupby('object_id').mjd.apply(lambda x: x.ptp())`\n\nwhere `ptp` stands for \"peak to peak\".",
          "votes": 1
        },
        {
          "id": 420071,
          "postDate": "2018-11-13T02:26:16.557Z",
          "content": "<p>This is way slower....  Apply is slow.  If you want a faster code, here is one:</p>\n\n<pre><code>   gr_mjd = train.groupby('object_id').mjd\n   train['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n</code></pre>\n\n<p>On my machine, your code takes 941 ms while mine takes 30 ms.  A 30x speedup ;)</p>\n\n<p>Unfortunately we cannot use ptp in transform.  It would be even faster.  Not sure why it isn't supported.</p>",
          "rawMarkdown": "This is way slower....  Apply is slow.  If you want a faster code, here is one:\n\n       gr_mjd = train.groupby('object_id').mjd\n       train['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n\nOn my machine, your code takes 941 ms while mine takes 30 ms.  A 30x speedup ;)\n\nUnfortunately we cannot use ptp in transform.  It would be even faster.  Not sure why it isn't supported.",
          "votes": 4
        },
        {
          "id": 434239,
          "postDate": "2018-12-06T04:54:25.890Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 410475,
      "postDate": "2018-10-26T05:33:23.733Z",
      "content": "<p>I see that some kernels use features based solely on when flux are measured, for instance <code>df['mjd_diff'] = df['mjd_max'] - df['mjd_min']</code></p>\n\n<p>Think of it.  How could the timing of measure influence the nature of the source we measure?  The flux we measure was emitted before it is measured.  There is no way the measure can influence the source.  </p>\n\n<p>Given there cannot be a causality effect, I am not using theses features.</p>\n\n<p>Edited.  Added 'solely' to make the point clearer.</p>",
      "rawMarkdown": "I see that some kernels use features based solely on when flux are measured, for instance `df['mjd_diff'] = df['mjd_max'] - df['mjd_min']`\n\nThink of it.  How could the timing of measure influence the nature of the source we measure?  The flux we measure was emitted before it is measured.  There is no way the measure can influence the source.  \n\nGiven there cannot be a causality effect, I am not using theses features.\n\nEdited.  Added 'solely' to make the point clearer.",
      "votes": 23
    },
    {
      "id": 410712,
      "postDate": "2018-10-26T14:10:49.273Z",
      "content": "<p>Don't forget that this isn't real data unfortunately - so they will be using mjds that theoretically observe the object - this could lead to information being leaked into the timings. You can do the analog of flux metrics for timings and you get ok results sadly.</p>",
      "rawMarkdown": "Don't forget that this isn't real data unfortunately - so they will be using mjds that theoretically observe the object - this could lead to information being leaked into the timings. You can do the analog of flux metrics for timings and you get ok results sadly.",
      "votes": 5,
      "replies": [
        {
          "id": 410719,
          "postDate": "2018-10-26T14:32:16.680Z",
          "content": "<blockquote>\n  <p>so they will be using mjds that theoretically observe the object</p>\n</blockquote>\n\n<p>Timings are related to the object localization indeed.</p>\n\n<p>I'd be surprised it this helps still.  I'll have a try at some point.</p>",
          "rawMarkdown": "&gt; so they will be using mjds that theoretically observe the object\n\nTimings are related to the object localization indeed.\n\nI'd be surprised it this helps still.  I'll have a try at some point.",
          "votes": 2
        }
      ]
    },
    {
      "id": 411382,
      "postDate": "2018-10-28T01:40:39.673Z",
      "content": "<p>I was about ask how this feature shared by Grzegorz is different from the feature used in the Olivier's kernel and I figured it.</p>\n\n<p>This feature make a huge boost in CV and LB scores.\nThanks Grzegorz for shared it.</p>",
      "rawMarkdown": "I was about ask how this feature shared by Grzegorz is different from the feature used in the Olivier's kernel and I figured it.\n\nThis feature make a huge boost in CV and LB scores.\nThanks Grzegorz for shared it.",
      "votes": 1
    },
    {
      "id": 410988,
      "postDate": "2018-10-27T04:25:15.460Z",
      "content": "<p>@CPMP, are you saying my feature engineering skills are bad ???</p>\n\n<p>LOL, I think I dropped this one a while back :)</p>",
      "rawMarkdown": "@CPMP, are you saying my feature engineering skills are bad ???\n\nLOL, I think I dropped this one a while back :)",
      "votes": 2,
      "replies": [
        {
          "id": 411002,
          "postDate": "2018-10-27T06:22:23.787Z",
          "content": "<blockquote>\n  <p>LOL, I think I dropped this one a while back :)</p>\n</blockquote>\n\n<p>I was sure you did drop it, but I was thinking of all who start with your public kernel. ;)</p>",
          "rawMarkdown": "&gt; LOL, I think I dropped this one a while back :)\n\nI was sure you did drop it, but I was thinking of all who start with your public kernel. ;)",
          "votes": 1
        },
        {
          "id": 411013,
          "postDate": "2018-10-27T07:07:28.057Z",
          "content": "<p>Yeah you're right I updated the public kernel accordingly, though I doubt anyone uses it ;-)</p>",
          "rawMarkdown": "Yeah you're right I updated the public kernel accordingly, though I doubt anyone uses it ;-)",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 410538,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-10-26T08:21:34.327000",
      "content": "<blockquote>\n  <p>I see that some kernels use features based on when flux are measured, ... I am not using theses features.</p>\n</blockquote>\n\n<p>@CPMP, I bet you will ;)</p>\n\n<pre><code>dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n</code></pre>\n\n<p>is a great feature to separate \"one event\" objects as supernovae from \"cyclic event\" objects as cepheids.</p>",
      "votes": 41,
      "replies": [
        {
          "id": 410573,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-26T09:33:17.593000",
          "content": "<p>This one depends on what you measure, I can use it.  Thanks for sharing.  I edited my post to make it clearer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411119,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-27T13:04:17.527000",
          "content": "<p>Grzegorz, your feature gave me a significant boost of about 0.07,  half of what I gained today,.  Thanks a lot for having shared it. I am sure others got a boost with it too!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 411155,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-27T14:18:16.533000",
          "content": "<p>That's what I've found in local CV. Hope it will translate to LB ;-)</p>\n\n<p>A big thank you for sharing with us <a href=\"/sionek\">@sionek</a></p>\n\n<p>UPDATE: I got a 0.1 LB improvement</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 415791,
          "author_name": "Michal Haltuf",
          "author_url": "",
          "post_date": "2018-11-05T17:45:11.250000",
          "content": "<p>Thank you, this feature really helps a lot to my model!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419065,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-11T07:02:24.720000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419072,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-11T07:29:02.383000",
          "content": "<p>That's using python table I think.  you can do something similar with pandas.groupby().  Something like:</p>\n\n<pre><code>train['mjd_min'] = train.groupby('object_id').mjd.transform('min')\ntrain['mjd_max'] = train.groupby('object_id').mjd.transform('max')\ntrain['mjd_diff']  = train['mjd_max']  - train['mjd_min'] \ndel train['mjd_min'] , train['mjd_max']\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 419989,
          "author_name": "Max Halford",
          "author_url": "",
          "post_date": "2018-11-12T22:09:12.880000",
          "content": "<p>Just for everyone's general knowledge:</p>\n\n<p><code>train['mjd_diff']  = train.groupby('object_id').mjd.apply(lambda x: x.ptp())</code></p>\n\n<p>where <code>ptp</code> stands for \"peak to peak\".</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 420071,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-13T02:26:16.557000",
          "content": "<p>This is way slower....  Apply is slow.  If you want a faster code, here is one:</p>\n\n<pre><code>   gr_mjd = train.groupby('object_id').mjd\n   train['mjd_diff']  = gr_mjd.transform('max') - gr_mjd.transform('min')\n</code></pre>\n\n<p>On my machine, your code takes 941 ms while mine takes 30 ms.  A 30x speedup ;)</p>\n\n<p>Unfortunately we cannot use ptp in transform.  It would be even faster.  Not sure why it isn't supported.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 434239,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-06T04:54:25.890000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 410712,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2018-10-26T14:10:49.273000",
      "content": "<p>Don't forget that this isn't real data unfortunately - so they will be using mjds that theoretically observe the object - this could lead to information being leaked into the timings. You can do the analog of flux metrics for timings and you get ok results sadly.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 410719,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-26T14:32:16.680000",
          "content": "<blockquote>\n  <p>so they will be using mjds that theoretically observe the object</p>\n</blockquote>\n\n<p>Timings are related to the object localization indeed.</p>\n\n<p>I'd be surprised it this helps still.  I'll have a try at some point.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 411382,
      "author_name": "João Pedro Peinado",
      "author_url": "",
      "post_date": "2018-10-28T01:40:39.673000",
      "content": "<p>I was about ask how this feature shared by Grzegorz is different from the feature used in the Olivier's kernel and I figured it.</p>\n\n<p>This feature make a huge boost in CV and LB scores.\nThanks Grzegorz for shared it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 410988,
      "author_name": "olivier",
      "author_url": "",
      "post_date": "2018-10-27T04:25:15.460000",
      "content": "<p>@CPMP, are you saying my feature engineering skills are bad ???</p>\n\n<p>LOL, I think I dropped this one a while back :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 411002,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-10-27T06:22:23.787000",
          "content": "<blockquote>\n  <p>LOL, I think I dropped this one a while back :)</p>\n</blockquote>\n\n<p>I was sure you did drop it, but I was thinking of all who start with your public kernel. ;)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 411013,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2018-10-27T07:07:28.057000",
          "content": "<p>Yeah you're right I updated the public kernel accordingly, though I doubt anyone uses it ;-)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "410538": "&gt; I see that some kernels use features based on when flux are measured, ... I am not using theses features.\n\n@CPMP, I bet you will ;)\n\n    dt[detected==1, mjd_diff:=max(mjd)-min(mjd), by=object_id]\n\nis a great feature to separate \"one event\" objects as supernovae from \"cyclic event\" objects as cepheids.",
    "410475": "I see that some kernels use features based solely on when flux are measured, for instance `df['mjd_diff'] = df['mjd_max'] - df['mjd_min']`\n\nThink of it.  How could the timing of measure influence the nature of the source we measure?  The flux we measure was emitted before it is measured.  There is no way the measure can influence the source.  \n\nGiven there cannot be a causality effect, I am not using theses features.\n\nEdited.  Added 'solely' to make the point clearer.",
    "410712": "Don't forget that this isn't real data unfortunately - so they will be using mjds that theoretically observe the object - this could lead to information being leaked into the timings. You can do the analog of flux metrics for timings and you get ok results sadly.",
    "411382": "I was about ask how this feature shared by Grzegorz is different from the feature used in the Olivier's kernel and I figured it.\n\nThis feature make a huge boost in CV and LB scores.\nThanks Grzegorz for shared it.",
    "410988": "@CPMP, are you saying my feature engineering skills are bad ???\n\nLOL, I think I dropped this one a while back :)"
  }
}