{
  "id": 27895,
  "title": "what's the important features did you generate?  ",
  "url": "/competitions/outbrain-click-prediction/discussion/27895",
  "author_name": "",
  "post_date": "2017-01-19T01:47:03.520Z",
  "votes": 6,
  "comment_count": 8,
  "views": 495,
  "content": "<p>Congratulations to the winners! So impressive 2 teams could break 0.7  and single model could also break 0.7. So what's the important features did you generate? Are they suitable for all models? </p>",
  "messages": [
    {
      "id": "157052",
      "postDate": "01/19/2017 01:47:03",
      "content": "<p>Congratulations to the winners! So impressive 2 teams could break 0.7  and single model could also break 0.7. So what's the important features did you generate? Are they suitable for all models? </p>",
      "rawMarkdown": "Congratulations to the winners! So impressive 2 teams could break 0.7  and single model could also break 0.7. So what's the important features did you generate? Are they suitable for all models?",
      "votes": null
    },
    {
      "id": "157105",
      "postDate": "01/19/2017 08:21:28",
      "content": "<p>I found several features that improved score significantly:</p>\n\n<ul>\n<li><p>first document that seen by uuid right after even and another features based on it (source, publisher etc). This improved my score for +.006</p></li>\n<li><p>another kind of leak: boolean value indicating whether ad will be shown for same user in the future (+.00160 to score)</p></li>\n<li><p>word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)</p></li>\n</ul>",
      "rawMarkdown": "I found several features that improved score significantly:\r\n\r\n- first document that seen by uuid right after even and another features based on it (source, publisher etc). This improved my score for +.006\r\n\r\n- another kind of leak: boolean value indicating whether ad will be shown for same user in the future (+.00160 to score)\r\n\r\n- word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)",
      "votes": null
    },
    {
      "id": "157128",
      "postDate": "01/19/2017 10:37:41",
      "content": "<p>[quote=rokh;157105]</p>\n\n<ul>\n<li>word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>very nice</p>",
      "rawMarkdown": "[quote=rokh;157105]\r\n\r\n- word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)\r\n\r\n[/quote]\r\n\r\nvery nice",
      "votes": null
    },
    {
      "id": "157175",
      "postDate": "01/19/2017 14:54:24",
      "content": "<p>@rokh Thanks for sharing. Could you please elaborate the first feature you generated? Is it a leak feature?  The next page(document_id) for one uuid visited is likely the landing page when he clicks the ad_id? </p>",
      "rawMarkdown": "rokh Thanks for sharing. Could you please elaborate the first feature you generated? Is it a leak feature?  The next page(document_id) for one uuid visited is likely the landing page when he clicks the ad_id?",
      "votes": null
    },
    {
      "id": "157216",
      "postDate": "01/19/2017 18:08:38",
      "content": "<p>@FengLi Yes, this is leak feature because this is information from the future, but it's not always document user is visiting by clicking the ad (though sometimes it is). It's important because it lets you know what user is looking for</p>",
      "rawMarkdown": "FengLi Yes, this is leak feature because this is information from the future, but it's not always document user is visiting by clicking the ad (though sometimes it is). It's important because it lets you know what user is looking for",
      "votes": null
    },
    {
      "id": "157229",
      "postDate": "01/19/2017 19:44:56",
      "content": "<p>[quote=rokh;157105]</p>\n\n<ul>\n<li>word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>this is awesome. How much improvement did it bring? Did it improve xgboost and FFM? I was thinking about that but I was too busy doing something else (that turned out to be useless :P)</p>",
      "rawMarkdown": "[quote=rokh;157105]\r\n\r\n- word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)\r\n\r\n[/quote]\r\n\r\nthis is awesome. How much improvement did it bring? Did it improve xgboost and FFM? I was thinking about that but I was too busy doing something else (that turned out to be useless :P)",
      "votes": null
    },
    {
      "id": "157313",
      "postDate": "01/20/2017 07:10:59",
      "content": "<p>@Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.</p>\n\n<p>I'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation</p>",
      "rawMarkdown": "Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.\r\n\r\nI'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation",
      "votes": null
    },
    {
      "id": "157314",
      "postDate": "01/20/2017 07:16:50",
      "content": "<p>I had userSeenThisAdTimes feature and similar. Assume that ffmlib implies the relationship, but haven't checked. First type of leak not required any post work at all</p>",
      "rawMarkdown": "I had userSeenThisAdTimes feature and similar. Assume that ffmlib implies the relationship, but haven't checked. First type of leak not required any post work at all",
      "votes": null
    },
    {
      "id": "157378",
      "postDate": "01/20/2017 16:30:29",
      "content": "<p>[quote=rokh;157313]</p>\n\n<p>@Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.</p>\n\n<p>I'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation</p>\n\n<p>[/quote]\nThanks!</p>\n\n<p>We have a similar feature if I understand yours correctly. I think it is just similar to first type of leak? I suspect that the timestamp sometimes can be off (e.g. latency) maybe a couple of hours even. So the doc right after ads shown might not be actually the actual doc users clicked. But including all future clicks would have the info. I did try clicks within 24 hours, which didn't add extra value. But your understanding makes sense too.</p>",
      "rawMarkdown": "[quote=rokh;157313]\r\n\r\n@Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.\r\n\r\nI'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation\r\n\r\n[/quote]\r\nThanks!\r\n\r\nWe have a similar feature if I understand yours correctly. I think it is just similar to first type of leak? I suspect that the timestamp sometimes can be off (e.g. latency) maybe a couple of hours even. So the doc right after ads shown might not be actually the actual doc users clicked. But including all future clicks would have the info. I did try clicks within 24 hours, which didn't add extra value. But your understanding makes sense too.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 157105,
      "author_name": "khanenko",
      "author_url": "",
      "post_date": "01/19/2017 08:21:28",
      "content": "<p>I found several features that improved score significantly:</p>\n\n<ul>\n<li><p>first document that seen by uuid right after even and another features based on it (source, publisher etc). This improved my score for +.006</p></li>\n<li><p>another kind of leak: boolean value indicating whether ad will be shown for same user in the future (+.00160 to score)</p></li>\n<li><p>word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 157128,
          "author_name": "cherednychenko",
          "author_url": "",
          "post_date": "01/19/2017 10:37:41",
          "content": "<p>[quote=rokh;157105]</p>\n\n<ul>\n<li>word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>very nice</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 157229,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "01/19/2017 19:44:56",
          "content": "<p>[quote=rokh;157105]</p>\n\n<ul>\n<li>word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)</li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>this is awesome. How much improvement did it bring? Did it improve xgboost and FFM? I was thinking about that but I was too busy doing something else (that turned out to be useless :P)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157175,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "01/19/2017 14:54:24",
      "content": "<p>@rokh Thanks for sharing. Could you please elaborate the first feature you generated? Is it a leak feature?  The next page(document_id) for one uuid visited is likely the landing page when he clicks the ad_id? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157216,
      "author_name": "khanenko",
      "author_url": "",
      "post_date": "01/19/2017 18:08:38",
      "content": "<p>@FengLi Yes, this is leak feature because this is information from the future, but it's not always document user is visiting by clicking the ad (though sometimes it is). It's important because it lets you know what user is looking for</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 157313,
      "author_name": "khanenko",
      "author_url": "",
      "post_date": "01/20/2017 07:10:59",
      "content": "<p>@Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.</p>\n\n<p>I'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation</p>",
      "votes": null,
      "replies": [
        {
          "id": 157378,
          "author_name": "xiaozhouwang",
          "author_url": "",
          "post_date": "01/20/2017 16:30:29",
          "content": "<p>[quote=rokh;157313]</p>\n\n<p>@Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.</p>\n\n<p>I'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation</p>\n\n<p>[/quote]\nThanks!</p>\n\n<p>We have a similar feature if I understand yours correctly. I think it is just similar to first type of leak? I suspect that the timestamp sometimes can be off (e.g. latency) maybe a couple of hours even. So the doc right after ads shown might not be actually the actual doc users clicked. But including all future clicks would have the info. I did try clicks within 24 hours, which didn't add extra value. But your understanding makes sense too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 157314,
      "author_name": "cherednychenko",
      "author_url": "",
      "post_date": "01/20/2017 07:16:50",
      "content": "<p>I had userSeenThisAdTimes feature and similar. Assume that ffmlib implies the relationship, but haven't checked. First type of leak not required any post work at all</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "157052": "Congratulations to the winners! So impressive 2 teams could break 0.7  and single model could also break 0.7. So what's the important features did you generate? Are they suitable for all models?",
    "157105": "I found several features that improved score significantly:\r\n\r\n- first document that seen by uuid right after even and another features based on it (source, publisher etc). This improved my score for +.006\r\n\r\n- another kind of leak: boolean value indicating whether ad will be shown for same user in the future (+.00160 to score)\r\n\r\n- word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)",
    "157128": "[quote=rokh;157105]\r\n\r\n- word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)\r\n\r\n[/quote]\r\n\r\nvery nice",
    "157175": "rokh Thanks for sharing. Could you please elaborate the first feature you generated? Is it a leak feature?  The next page(document_id) for one uuid visited is likely the landing page when he clicks the ad_id?",
    "157216": "FengLi Yes, this is leak feature because this is information from the future, but it's not always document user is visiting by clicking the ad (though sometimes it is). It's important because it lets you know what user is looking for",
    "157229": "[quote=rokh;157105]\r\n\r\n- word embeddings for documents based on page views (document_id is word, while sentences - consecutive page views of same user)\r\n\r\n[/quote]\r\n\r\nthis is awesome. How much improvement did it bring? Did it improve xgboost and FFM? I was thinking about that but I was too busy doing something else (that turned out to be useless :P)",
    "157313": "Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.\r\n\r\nI'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation",
    "157314": "I had userSeenThisAdTimes feature and similar. Assume that ffmlib implies the relationship, but haven't checked. First type of leak not required any post work at all",
    "157378": "[quote=rokh;157313]\r\n\r\n@Little Boat Well, I'm not sure. I added features based on w2v at the very begining of competition and did not spend much time to tune it. Maybe 0.0005-0.001 or even less, so it's  probably not so important but still interesting to try.\r\n\r\nI'm curious if someone from top3 used second leak feature I mentioned above because it added 0.0016 to my score several days before final deadline. It seems that outbrain rarely show documents in ads that user had visited before. So it means if same ad was shown 2 times for same user it likely was not clicked in first observation\r\n\r\n[/quote]\r\nThanks!\r\n\r\nWe have a similar feature if I understand yours correctly. I think it is just similar to first type of leak? I suspect that the timestamp sometimes can be off (e.g. latency) maybe a couple of hours even. So the doc right after ads shown might not be actually the actual doc users clicked. But including all future clicks would have the info. I did try clicks within 24 hours, which didn't add extra value. But your understanding makes sense too."
  },
  "source": "meta"
}