{
  "id": 55630,
  "title": "Include both Region and City params or just one?",
  "url": "/competitions/avito-demand-prediction/discussion/55630",
  "author_name": "",
  "post_date": "2018-04-30T02:54:18.470029400Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm new to machine learning, but I sure am curious.  I was wondering if I should include both <code>region</code> and <code>city</code> from the data files in my decision tree model.  The correlation between them is extremely high, as city can always determine the region.  Is this good logic?</p>",
  "messages": [
    {
      "id": "320849",
      "postDate": "04/30/2018 02:54:18",
      "content": "<p>I'm new to machine learning, but I sure am curious.  I was wondering if I should include both <code>region</code> and <code>city</code> from the data files in my decision tree model.  The correlation between them is extremely high, as city can always determine the region.  Is this good logic?</p>",
      "rawMarkdown": "I'm new to machine learning, but I sure am curious.  I was wondering if I should include both `region` and `city` from the data files in my decision tree model.  The correlation between them is extremely high, as city can always determine the region.  Is this good logic?",
      "votes": null
    },
    {
      "id": "320865",
      "postDate": "04/30/2018 03:57:43",
      "content": "<p>There's some reason to think to exclude one, since they have high correlation and some models don't like that. But there's also some reason to keep both, since region is more information about the higher-level structure of the cities and cities is more information about the lower-level structure of region.</p>\n\n<p>Why not try both and see which has a better cross-validation score? :)</p>",
      "rawMarkdown": "There's some reason to think to exclude one, since they have high correlation and some models don't like that. But there's also some reason to keep both, since region is more information about the higher-level structure of the cities and cities is more information about the lower-level structure of region.\n\nWhy not try both and see which has a better cross-validation score? :)",
      "votes": null
    },
    {
      "id": "321590",
      "postDate": "05/01/2018 16:05:28",
      "content": "<p>also, try to combine the cat features togeter such as:\ndf['new_feature']  = df['feature_x'] +'  '+ df['feature_y']</p>",
      "rawMarkdown": "also, try to combine the cat features togeter such as:\ndf['new_feature']  = df['feature_x'] +'  '+ df['feature_y']",
      "votes": null
    },
    {
      "id": "321591",
      "postDate": "05/01/2018 16:07:38",
      "content": "<p>Combining <code>region</code> and <code>city</code> would have the same effectiveness as just <code>city</code> on its own (unless there is something trivial about the model such as overfit).  I will definitely try that for other things though.</p>",
      "rawMarkdown": "Combining `region` and `city` would have the same effectiveness as just `city` on its own (unless there is something trivial about the model such as overfit).  I will definitely try that for other things though.",
      "votes": null
    },
    {
      "id": "321751",
      "postDate": "05/01/2018 20:53:28",
      "content": "<blockquote>\n  <p><strong>Matthew Anderson wrote</strong></p>\n  \n  <blockquote>\n    <p>Combining <code>region</code> and <code>city</code> would have the same effectiveness as just <code>city</code> on its own</p>\n  </blockquote>\n</blockquote>\n\n<p>That isn't necessarily true, the same city name may be used in multiple regions. The city of \"Октябрьский\", for example, is in 7 different regions in the training set. This is similar to having cities named \"Springfield\" in 41 different states in the US. Combining the two features will help distinguish between them.</p>",
      "rawMarkdown": "&gt; **Matthew Anderson wrote**\n&gt; \n&gt; &gt; Combining `region` and `city` would have the same effectiveness as just `city` on its own\n\nThat isn't necessarily true, the same city name may be used in multiple regions. The city of \"Октябрьский\", for example, is in 7 different regions in the training set. This is similar to having cities named \"Springfield\" in 41 different states in the US. Combining the two features will help distinguish between them.",
      "votes": null
    },
    {
      "id": "321782",
      "postDate": "05/01/2018 22:05:26",
      "content": "<p>Thank you Brandon!  It did not occur to me that such a thing could be happening.  I will try that in my FE.</p>",
      "rawMarkdown": "Thank you Brandon!  It did not occur to me that such a thing could be happening.  I will try that in my FE.",
      "votes": null
    },
    {
      "id": "321807",
      "postDate": "05/02/2018 00:03:33",
      "content": "<p>Yeah, that's a good point, Branden! I didn't think of that at all either. Have my upvote!</p>",
      "rawMarkdown": "Yeah, that's a good point, Branden! I didn't think of that at all either. Have my upvote!",
      "votes": null
    },
    {
      "id": "323572",
      "postDate": "05/05/2018 15:00:35",
      "content": "<p>Nice!</p>",
      "rawMarkdown": "Nice!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 320865,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "04/30/2018 03:57:43",
      "content": "<p>There's some reason to think to exclude one, since they have high correlation and some models don't like that. But there's also some reason to keep both, since region is more information about the higher-level structure of the cities and cities is more information about the lower-level structure of region.</p>\n\n<p>Why not try both and see which has a better cross-validation score? :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321590,
      "author_name": "mtinti",
      "author_url": "",
      "post_date": "05/01/2018 16:05:28",
      "content": "<p>also, try to combine the cat features togeter such as:\ndf['new_feature']  = df['feature_x'] +'  '+ df['feature_y']</p>",
      "votes": null,
      "replies": [
        {
          "id": 321591,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/01/2018 16:07:38",
          "content": "<p>Combining <code>region</code> and <code>city</code> would have the same effectiveness as just <code>city</code> on its own (unless there is something trivial about the model such as overfit).  I will definitely try that for other things though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321751,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "05/01/2018 20:53:28",
          "content": "<blockquote>\n  <p><strong>Matthew Anderson wrote</strong></p>\n  \n  <blockquote>\n    <p>Combining <code>region</code> and <code>city</code> would have the same effectiveness as just <code>city</code> on its own</p>\n  </blockquote>\n</blockquote>\n\n<p>That isn't necessarily true, the same city name may be used in multiple regions. The city of \"Октябрьский\", for example, is in 7 different regions in the training set. This is similar to having cities named \"Springfield\" in 41 different states in the US. Combining the two features will help distinguish between them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321782,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/01/2018 22:05:26",
          "content": "<p>Thank you Brandon!  It did not occur to me that such a thing could be happening.  I will try that in my FE.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321807,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "05/02/2018 00:03:33",
          "content": "<p>Yeah, that's a good point, Branden! I didn't think of that at all either. Have my upvote!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323572,
          "author_name": "stonedl",
          "author_url": "",
          "post_date": "05/05/2018 15:00:35",
          "content": "<p>Nice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "320849": "I'm new to machine learning, but I sure am curious.  I was wondering if I should include both `region` and `city` from the data files in my decision tree model.  The correlation between them is extremely high, as city can always determine the region.  Is this good logic?",
    "320865": "There's some reason to think to exclude one, since they have high correlation and some models don't like that. But there's also some reason to keep both, since region is more information about the higher-level structure of the cities and cities is more information about the lower-level structure of region.\n\nWhy not try both and see which has a better cross-validation score? :)",
    "321590": "also, try to combine the cat features togeter such as:\ndf['new_feature']  = df['feature_x'] +'  '+ df['feature_y']",
    "321591": "Combining `region` and `city` would have the same effectiveness as just `city` on its own (unless there is something trivial about the model such as overfit).  I will definitely try that for other things though.",
    "321751": "&gt; **Matthew Anderson wrote**\n&gt; \n&gt; &gt; Combining `region` and `city` would have the same effectiveness as just `city` on its own\n\nThat isn't necessarily true, the same city name may be used in multiple regions. The city of \"Октябрьский\", for example, is in 7 different regions in the training set. This is similar to having cities named \"Springfield\" in 41 different states in the US. Combining the two features will help distinguish between them.",
    "321782": "Thank you Brandon!  It did not occur to me that such a thing could be happening.  I will try that in my FE.",
    "321807": "Yeah, that's a good point, Branden! I didn't think of that at all either. Have my upvote!",
    "323572": "Nice!"
  },
  "source": "meta"
}