{
  "id": 21821,
  "title": "Question to dig",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/21821",
  "author_name": "",
  "post_date": "2016-06-21T09:30:37.853Z",
  "votes": null,
  "comment_count": 7,
  "views": 701,
  "content": "<p>Hi</p>\n\n<p>I got a query for Kaggle, and Sponsors specifically, and may be something to dig, for fellow Kagglers. And this might be applicable for other competitions too. Lets say some features which are purely ID e.g (category_ID),  in their pure form, help in improving predictions by some fraction and therefore allows me or a team to get up there and win a competition or even in the top 3 or 5 etc. Point is, what would that give to the Sponsors ? or the Community ?. My guess is not much because IDs can change tomorrow and they have no relevance what soever (unless there is a leak) with the data.</p>\n\n<p>So - should we use such features ? Or may be - The rules should say that - &quot;you cannot use IDs as features&quot;. </p>\n\n<p>What do you think ?</p>\n\n<p>Regds</p>",
  "messages": [
    {
      "id": "124687",
      "postDate": "06/21/2016 09:30:37",
      "content": "<p>Hi</p>\n\n<p>I got a query for Kaggle, and Sponsors specifically, and may be something to dig, for fellow Kagglers. And this might be applicable for other competitions too. Lets say some features which are purely ID e.g (category_ID),  in their pure form, help in improving predictions by some fraction and therefore allows me or a team to get up there and win a competition or even in the top 3 or 5 etc. Point is, what would that give to the Sponsors ? or the Community ?. My guess is not much because IDs can change tomorrow and they have no relevance what soever (unless there is a leak) with the data.</p>\n\n<p>So - should we use such features ? Or may be - The rules should say that - &quot;you cannot use IDs as features&quot;. </p>\n\n<p>What do you think ?</p>\n\n<p>Regds</p>",
      "rawMarkdown": "Hi\r\n\r\nI got a query for Kaggle, and Sponsors specifically, and may be something to dig, for fellow Kagglers. And this might be applicable for other competitions too. Lets say some features which are purely ID e.g (category_ID),  in their pure form, help in improving predictions by some fraction and therefore allows me or a team to get up there and win a competition or even in the top 3 or 5 etc. Point is, what would that give to the Sponsors ? or the Community ?. My guess is not much because IDs can change tomorrow and they have no relevance what soever (unless there is a leak) with the data.\r\n\r\nSo - should we use such features ? Or may be - The rules should say that - \"you cannot use IDs as features\". \r\n\r\nWhat do you think ?\r\n\r\nRegds",
      "votes": null
    },
    {
      "id": "124689",
      "postDate": "06/21/2016 09:58:22",
      "content": "<p>I think category id has meaning, and if the id changes, the meaning should stay the same</p>",
      "rawMarkdown": "I think category id has meaning, and if the id changes, the meaning should stay the same",
      "votes": null
    },
    {
      "id": "124690",
      "postDate": "06/21/2016 10:03:18",
      "content": "<p>[quote=ololo;124689]</p>\n\n<p>I think category id has meaning, and if the id changes, the meaning should stay the same</p>\n\n<p>[/quote]\nOlolo. I am not talking about handling different categories differently. I am talking about using &quot;their value&quot; as is. The value of a category (as we see between 9 and 117) cannot hold a meaning in itself (in regression or in trees). I am asking peoples view on submitting better performing models (some how, specific to this data)  which use the values as is.</p>",
      "rawMarkdown": "[quote=ololo;124689]\r\n\r\nI think category id has meaning, and if the id changes, the meaning should stay the same\r\n\r\n[/quote]\r\nOlolo. I am not talking about handling different categories differently. I am talking about using \"their value\" as is. The value of a category (as we see between 9 and 117) cannot hold a meaning in itself (in regression or in trees). I am asking peoples view on submitting better performing models (some how, specific to this data)  which use the values as is.",
      "votes": null
    },
    {
      "id": "124692",
      "postDate": "06/21/2016 10:16:57",
      "content": "<p>[quote=Run2;124690]\nhandling different categories differently.\n[/quote]</p>\n\n<p>I think this is exactly what you are doing when including category_id (as is) to your tree-based model. But for linear models it obviously won't work. </p>",
      "rawMarkdown": "[quote=Run2;124690]\r\nhandling different categories differently.\r\n[/quote]\r\n\r\nI think this is exactly what you are doing when including category_id (as is) to your tree-based model. But for linear models it obviously won't work.",
      "votes": null
    },
    {
      "id": "124693",
      "postDate": "06/21/2016 10:26:48",
      "content": "<p>[quote=ololo;124692]</p>\n\n<p>[quote=Run2;124690]\nhandling different categories differently.\n[/quote]</p>\n\n<p>I think this is exactly what you are doing when including category ids to your tree-based model. But for linear models it obviously won't work. </p>\n\n<p>[/quote]\nEven for tree based models - you should not use the values as is - you should use one hot features. If you use the values in tree based model again the model will fail if the values change.</p>\n\n<p>And my query was specifically on that &quot;won't work&quot; park. Just say it works and some how the predictions are better using the bare values - and I win ( dreaming :-) ). Is that a valid win ?</p>",
      "rawMarkdown": "[quote=ololo;124692]\r\n\r\n[quote=Run2;124690]\r\nhandling different categories differently.\r\n[/quote]\r\n\r\nI think this is exactly what you are doing when including category ids to your tree-based model. But for linear models it obviously won't work. \r\n\r\n[/quote]\r\nEven for tree based models - you should not use the values as is - you should use one hot features. If you use the values in tree based model again the model will fail if the values change.\r\n\r\nAnd my query was specifically on that \"won't work\" park. Just say it works and some how the predictions are better using the bare values - and I win ( dreaming :-) ). Is that a valid win ?",
      "votes": null
    },
    {
      "id": "125884",
      "postDate": "07/04/2016 02:06:01",
      "content": "<p>So, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?</p>",
      "rawMarkdown": "So, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?",
      "votes": null
    },
    {
      "id": "125888",
      "postDate": "07/04/2016 03:14:05",
      "content": "<p>[quote=TheKai;125884]</p>\n\n<p>So, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?</p>\n\n<p>[/quote]</p>\n\n<p>No mistake, categs are equal.</p>",
      "rawMarkdown": "[quote=TheKai;125884]\r\n\r\nSo, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?\r\n\r\n[/quote]\r\n\r\nNo mistake, categs are equal.",
      "votes": null
    },
    {
      "id": "125890",
      "postDate": "07/04/2016 03:35:18",
      "content": "<p>Ah, thanks Snow Dog!</p>",
      "rawMarkdown": "Ah, thanks Snow Dog!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 124689,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "06/21/2016 09:58:22",
      "content": "<p>I think category id has meaning, and if the id changes, the meaning should stay the same</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124690,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "06/21/2016 10:03:18",
      "content": "<p>[quote=ololo;124689]</p>\n\n<p>I think category id has meaning, and if the id changes, the meaning should stay the same</p>\n\n<p>[/quote]\nOlolo. I am not talking about handling different categories differently. I am talking about using &quot;their value&quot; as is. The value of a category (as we see between 9 and 117) cannot hold a meaning in itself (in regression or in trees). I am asking peoples view on submitting better performing models (some how, specific to this data)  which use the values as is.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124692,
      "author_name": "agrigorev",
      "author_url": "",
      "post_date": "06/21/2016 10:16:57",
      "content": "<p>[quote=Run2;124690]\nhandling different categories differently.\n[/quote]</p>\n\n<p>I think this is exactly what you are doing when including category_id (as is) to your tree-based model. But for linear models it obviously won't work. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124693,
      "author_name": "rightfit",
      "author_url": "",
      "post_date": "06/21/2016 10:26:48",
      "content": "<p>[quote=ololo;124692]</p>\n\n<p>[quote=Run2;124690]\nhandling different categories differently.\n[/quote]</p>\n\n<p>I think this is exactly what you are doing when including category ids to your tree-based model. But for linear models it obviously won't work. </p>\n\n<p>[/quote]\nEven for tree based models - you should not use the values as is - you should use one hot features. If you use the values in tree based model again the model will fail if the values change.</p>\n\n<p>And my query was specifically on that &quot;won't work&quot; park. Just say it works and some how the predictions are better using the bare values - and I win ( dreaming :-) ). Is that a valid win ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125884,
      "author_name": "larsenk",
      "author_url": "",
      "post_date": "07/04/2016 02:06:01",
      "content": "<p>So, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125888,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "07/04/2016 03:14:05",
      "content": "<p>[quote=TheKai;125884]</p>\n\n<p>So, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?</p>\n\n<p>[/quote]</p>\n\n<p>No mistake, categs are equal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125890,
      "author_name": "larsenk",
      "author_url": "",
      "post_date": "07/04/2016 03:35:18",
      "content": "<p>Ah, thanks Snow Dog!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "124687": "Hi\r\n\r\nI got a query for Kaggle, and Sponsors specifically, and may be something to dig, for fellow Kagglers. And this might be applicable for other competitions too. Lets say some features which are purely ID e.g (category_ID),  in their pure form, help in improving predictions by some fraction and therefore allows me or a team to get up there and win a competition or even in the top 3 or 5 etc. Point is, what would that give to the Sponsors ? or the Community ?. My guess is not much because IDs can change tomorrow and they have no relevance what soever (unless there is a leak) with the data.\r\n\r\nSo - should we use such features ? Or may be - The rules should say that - \"you cannot use IDs as features\". \r\n\r\nWhat do you think ?\r\n\r\nRegds",
    "124689": "I think category id has meaning, and if the id changes, the meaning should stay the same",
    "124690": "[quote=ololo;124689]\r\n\r\nI think category id has meaning, and if the id changes, the meaning should stay the same\r\n\r\n[/quote]\r\nOlolo. I am not talking about handling different categories differently. I am talking about using \"their value\" as is. The value of a category (as we see between 9 and 117) cannot hold a meaning in itself (in regression or in trees). I am asking peoples view on submitting better performing models (some how, specific to this data)  which use the values as is.",
    "124692": "[quote=Run2;124690]\r\nhandling different categories differently.\r\n[/quote]\r\n\r\nI think this is exactly what you are doing when including category_id (as is) to your tree-based model. But for linear models it obviously won't work.",
    "124693": "[quote=ololo;124692]\r\n\r\n[quote=Run2;124690]\r\nhandling different categories differently.\r\n[/quote]\r\n\r\nI think this is exactly what you are doing when including category ids to your tree-based model. But for linear models it obviously won't work. \r\n\r\n[/quote]\r\nEven for tree based models - you should not use the values as is - you should use one hot features. If you use the values in tree based model again the model will fail if the values change.\r\n\r\nAnd my query was specifically on that \"won't work\" park. Just say it works and some how the predictions are better using the bare values - and I win ( dreaming :-) ). Is that a valid win ?",
    "125884": "So, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?",
    "125888": "[quote=TheKai;125884]\r\n\r\nSo, I'm clearly doing something wrong, but it seems like every pair in the training set refers to the same category. Could someone confirm that I'm making some mistake in my joins leading to this incorrect (if I go by the above posts) assumption?\r\n\r\n[/quote]\r\n\r\nNo mistake, categs are equal.",
    "125890": "Ah, thanks Snow Dog!"
  },
  "source": "meta"
}