{
  "id": 53815,
  "title": "Newbie question: is all this hassle about the IP encoding normal for Kaggle?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53815",
  "author_name": "",
  "post_date": "2018-04-05T13:51:17.344998800Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I feel like much of this competition has been trying to figure out how it was set up, with less emphasis on interesting analysis / feature engineering / model selection / etc. Granted, figuring out exactly what TalkingData did is somewhat interesting, but I can't help but feel that it distracts from actually solving the problem. Is all this normal for Kaggle competitions, or is this something of an outlier? If it is normal, do you old-timers think it adds to the experience, or detracts?</p>",
  "messages": [
    {
      "id": "309510",
      "postDate": "04/05/2018 13:51:17",
      "content": "<p>I feel like much of this competition has been trying to figure out how it was set up, with less emphasis on interesting analysis / feature engineering / model selection / etc. Granted, figuring out exactly what TalkingData did is somewhat interesting, but I can't help but feel that it distracts from actually solving the problem. Is all this normal for Kaggle competitions, or is this something of an outlier? If it is normal, do you old-timers think it adds to the experience, or detracts?</p>",
      "rawMarkdown": "I feel like much of this competition has been trying to figure out how it was set up, with less emphasis on interesting analysis / feature engineering / model selection / etc. Granted, figuring out exactly what TalkingData did is somewhat interesting, but I can't help but feel that it distracts from actually solving the problem. Is all this normal for Kaggle competitions, or is this something of an outlier? If it is normal, do you old-timers think it adds to the experience, or detracts?",
      "votes": null
    },
    {
      "id": "309528",
      "postDate": "04/05/2018 14:31:07",
      "content": "<p>how do you do analysis/fe/model selection without understanding you data? </p>",
      "rawMarkdown": "how do you do analysis/fe/model selection without understanding you data?",
      "votes": null
    },
    {
      "id": "309541",
      "postDate": "04/05/2018 14:48:35",
      "content": "<p>In case if competition is set up correctly, then yes - there is not much sense to know, How exactly it was set up and How exactly data were anonimized. </p>\n\n<p>But from my experience - in case if setup has some flaws, that could result in a leak. And in this case it is really important to know that as soon as possible. So for me it is already quite standard practice to make some trivial leak-tests as one of first things for every new competition. And if some variable (like IP here) looks somehow \"not usual\" it's definitely worth to explore that deeper just to be sure.</p>",
      "rawMarkdown": "In case if competition is set up correctly, then yes - there is not much sense to know, How exactly it was set up and How exactly data were anonimized. \n\nBut from my experience - in case if setup has some flaws, that could result in a leak. And in this case it is really important to know that as soon as possible. So for me it is already quite standard practice to make some trivial leak-tests as one of first things for every new competition. And if some variable (like IP here) looks somehow \"not usual\" it's definitely worth to explore that deeper just to be sure.",
      "votes": null
    },
    {
      "id": "309546",
      "postDate": "04/05/2018 14:54:50",
      "content": "<p>That's wonderful for natural features of the data, of course! But I mean more that this whole IP thing seems to be a manufactured hassle. In the real world, test data mirrors training data because we design it that way. Here, there has been a lot of time spent on reverse engineering IP encoding when, in practice, we wouldn't have to do that because we would have engineered it ourselves in the first place. Does that make sense?</p>",
      "rawMarkdown": "That's wonderful for natural features of the data, of course! But I mean more that this whole IP thing seems to be a manufactured hassle. In the real world, test data mirrors training data because we design it that way. Here, there has been a lot of time spent on reverse engineering IP encoding when, in practice, we wouldn't have to do that because we would have engineered it ourselves in the first place. Does that make sense?",
      "votes": null
    },
    {
      "id": "309548",
      "postDate": "04/05/2018 14:57:57",
      "content": "<p>This makes sense. Thanks!</p>",
      "rawMarkdown": "This makes sense. Thanks!",
      "votes": null
    },
    {
      "id": "309667",
      "postDate": "04/05/2018 19:17:58",
      "content": "<blockquote>\n  <p>If it is normal, do you old-timers think it adds to the experience, or detracts?</p>\n</blockquote>\n\n<p>You probably should not be expecting many responses when asking old-timers specifically to speak up ;-)</p>",
      "rawMarkdown": "&gt; If it is normal, do you old-timers think it adds to the experience, or detracts?\n\nYou probably should not be expecting many responses when asking old-timers specifically to speak up ;-)",
      "votes": null
    },
    {
      "id": "309674",
      "postDate": "04/05/2018 19:26:43",
      "content": "<p>Haha I was merely trying to be deferential to my betters. :)</p>",
      "rawMarkdown": "Haha I was merely trying to be deferential to my betters. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 309528,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "04/05/2018 14:31:07",
      "content": "<p>how do you do analysis/fe/model selection without understanding you data? </p>",
      "votes": null,
      "replies": [
        {
          "id": 309546,
          "author_name": "chadwgardner",
          "author_url": "",
          "post_date": "04/05/2018 14:54:50",
          "content": "<p>That's wonderful for natural features of the data, of course! But I mean more that this whole IP thing seems to be a manufactured hassle. In the real world, test data mirrors training data because we design it that way. Here, there has been a lot of time spent on reverse engineering IP encoding when, in practice, we wouldn't have to do that because we would have engineered it ourselves in the first place. Does that make sense?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309541,
      "author_name": "alijs1",
      "author_url": "",
      "post_date": "04/05/2018 14:48:35",
      "content": "<p>In case if competition is set up correctly, then yes - there is not much sense to know, How exactly it was set up and How exactly data were anonimized. </p>\n\n<p>But from my experience - in case if setup has some flaws, that could result in a leak. And in this case it is really important to know that as soon as possible. So for me it is already quite standard practice to make some trivial leak-tests as one of first things for every new competition. And if some variable (like IP here) looks somehow \"not usual\" it's definitely worth to explore that deeper just to be sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 309548,
          "author_name": "chadwgardner",
          "author_url": "",
          "post_date": "04/05/2018 14:57:57",
          "content": "<p>This makes sense. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309667,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/05/2018 19:17:58",
      "content": "<blockquote>\n  <p>If it is normal, do you old-timers think it adds to the experience, or detracts?</p>\n</blockquote>\n\n<p>You probably should not be expecting many responses when asking old-timers specifically to speak up ;-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 309674,
          "author_name": "chadwgardner",
          "author_url": "",
          "post_date": "04/05/2018 19:26:43",
          "content": "<p>Haha I was merely trying to be deferential to my betters. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "309510": "I feel like much of this competition has been trying to figure out how it was set up, with less emphasis on interesting analysis / feature engineering / model selection / etc. Granted, figuring out exactly what TalkingData did is somewhat interesting, but I can't help but feel that it distracts from actually solving the problem. Is all this normal for Kaggle competitions, or is this something of an outlier? If it is normal, do you old-timers think it adds to the experience, or detracts?",
    "309528": "how do you do analysis/fe/model selection without understanding you data?",
    "309541": "In case if competition is set up correctly, then yes - there is not much sense to know, How exactly it was set up and How exactly data were anonimized. \n\nBut from my experience - in case if setup has some flaws, that could result in a leak. And in this case it is really important to know that as soon as possible. So for me it is already quite standard practice to make some trivial leak-tests as one of first things for every new competition. And if some variable (like IP here) looks somehow \"not usual\" it's definitely worth to explore that deeper just to be sure.",
    "309546": "That's wonderful for natural features of the data, of course! But I mean more that this whole IP thing seems to be a manufactured hassle. In the real world, test data mirrors training data because we design it that way. Here, there has been a lot of time spent on reverse engineering IP encoding when, in practice, we wouldn't have to do that because we would have engineered it ourselves in the first place. Does that make sense?",
    "309548": "This makes sense. Thanks!",
    "309667": "&gt; If it is normal, do you old-timers think it adds to the experience, or detracts?\n\nYou probably should not be expecting many responses when asking old-timers specifically to speak up ;-)",
    "309674": "Haha I was merely trying to be deferential to my betters. :)"
  },
  "source": "meta"
}