{
  "id": 15615,
  "title": "Any user parameter blows up the model?",
  "url": "/competitions/avito-context-ad-clicks/discussion/15615",
  "author_name": "",
  "post_date": "2015-07-29T11:39:54.130Z",
  "votes": null,
  "comment_count": 1,
  "views": 852,
  "content": "<p>Congrats to those who won. </p>\n\n<p>So I was cursed in this competition for some reason; my validation score blew up when I used anything related to a user. This is no longer relevant for the competition, but I'd really like to know what I did wrong for such a basic issue.</p>\n\n<p>I had training data lines like this:</p>\n\n<pre><code>-1 1 |u u1:570979 u2:4159 u3:63 u4:3 u5:3566 u6:0 u7:53 |a a8:16904117 a9:7 a10:38 a11:3 a12:41511.0 |s s13:131 s14:190 s15:-1 s16:-1 |b b18:2 b19:26 |c c21:26 c22:18 c23:1 |h h24:0.002913\n</code></pre>\n\n<p>|u are things from the userinfo table, |a adsinfo data, |s day/hour/city/region, |b ad statistics, |c user statistics, |h histctr.</p>\n\n<p>The user statistics are not directly statistically related; c21 is a representation of the click through rate adjusted for users that have really small numbers of clicks using an algorithm similar to a wilson score. Low number of clicks keep the user as &quot;average&quot; and anything &gt;10 already distinguishes this user from others. c22 is simply the number of days the user prowls the site. c23 is only used to figure out if a user appears in the trainstream or not.</p>\n\n<p>I could selectively disable any feature. Without |c features (user), I got cv results that approximated what I got on the LB, but huge differences when any user feature got activated. I only used FTRL. Example:</p>\n\n<p>without user features c21 and c22:\ntraining: 0.035, cv: 0.046, LB: 0.045\nwith user features c21 and c22:\ntraining: 0.029, cv: 0.038, LB: 0.061</p>\n\n<p>Training data was sorted by date, then hour. The exact same code for training, validation and test got used each time.</p>\n\n<p>From the entire training set, I collected impressions for every user and stored them separately for every user. Then when the day was finished, I wrote the last two impressions for every user to validation and everything else to the remainder of the training set.</p>\n\n<p>Does anyone spot an obvious flaw in this strategy?\nHave you been able to use any user features directly without suffering a similar issue?  What did you change to the model and learning to be able to use them?</p>\n\n<p>EDIT: changed b18/19 -&gt; c21/22</p>",
  "messages": [
    {
      "id": "87407",
      "postDate": "07/29/2015 11:39:54",
      "content": "<p>Congrats to those who won. </p>\n\n<p>So I was cursed in this competition for some reason; my validation score blew up when I used anything related to a user. This is no longer relevant for the competition, but I'd really like to know what I did wrong for such a basic issue.</p>\n\n<p>I had training data lines like this:</p>\n\n<pre><code>-1 1 |u u1:570979 u2:4159 u3:63 u4:3 u5:3566 u6:0 u7:53 |a a8:16904117 a9:7 a10:38 a11:3 a12:41511.0 |s s13:131 s14:190 s15:-1 s16:-1 |b b18:2 b19:26 |c c21:26 c22:18 c23:1 |h h24:0.002913\n</code></pre>\n\n<p>|u are things from the userinfo table, |a adsinfo data, |s day/hour/city/region, |b ad statistics, |c user statistics, |h histctr.</p>\n\n<p>The user statistics are not directly statistically related; c21 is a representation of the click through rate adjusted for users that have really small numbers of clicks using an algorithm similar to a wilson score. Low number of clicks keep the user as &quot;average&quot; and anything &gt;10 already distinguishes this user from others. c22 is simply the number of days the user prowls the site. c23 is only used to figure out if a user appears in the trainstream or not.</p>\n\n<p>I could selectively disable any feature. Without |c features (user), I got cv results that approximated what I got on the LB, but huge differences when any user feature got activated. I only used FTRL. Example:</p>\n\n<p>without user features c21 and c22:\ntraining: 0.035, cv: 0.046, LB: 0.045\nwith user features c21 and c22:\ntraining: 0.029, cv: 0.038, LB: 0.061</p>\n\n<p>Training data was sorted by date, then hour. The exact same code for training, validation and test got used each time.</p>\n\n<p>From the entire training set, I collected impressions for every user and stored them separately for every user. Then when the day was finished, I wrote the last two impressions for every user to validation and everything else to the remainder of the training set.</p>\n\n<p>Does anyone spot an obvious flaw in this strategy?\nHave you been able to use any user features directly without suffering a similar issue?  What did you change to the model and learning to be able to use them?</p>\n\n<p>EDIT: changed b18/19 -&gt; c21/22</p>",
      "rawMarkdown": "Congrats to those who won. \r\n\r\nSo I was cursed in this competition for some reason; my validation score blew up when I used anything related to a user. This is no longer relevant for the competition, but I'd really like to know what I did wrong for such a basic issue.\r\n\r\nI had training data lines like this:\r\n\r\n    -1 1 |u u1:570979 u2:4159 u3:63 u4:3 u5:3566 u6:0 u7:53 |a a8:16904117 a9:7 a10:38 a11:3 a12:41511.0 |s s13:131 s14:190 s15:-1 s16:-1 |b b18:2 b19:26 |c c21:26 c22:18 c23:1 |h h24:0.002913\r\n\r\n|u are things from the userinfo table, |a adsinfo data, |s day/hour/city/region, |b ad statistics, |c user statistics, |h histctr.\r\n\r\nThe user statistics are not directly statistically related; c21 is a representation of the click through rate adjusted for users that have really small numbers of clicks using an algorithm similar to a wilson score. Low number of clicks keep the user as \"average\" and anything >10 already distinguishes this user from others. c22 is simply the number of days the user prowls the site. c23 is only used to figure out if a user appears in the trainstream or not.\r\n\r\nI could selectively disable any feature. Without |c features (user), I got cv results that approximated what I got on the LB, but huge differences when any user feature got activated. I only used FTRL. Example:\r\n\r\nwithout user features c21 and c22:\r\ntraining: 0.035, cv: 0.046, LB: 0.045\r\nwith user features c21 and c22:\r\ntraining: 0.029, cv: 0.038, LB: 0.061\r\n\r\nTraining data was sorted by date, then hour. The exact same code for training, validation and test got used each time.\r\n\r\nFrom the entire training set, I collected impressions for every user and stored them separately for every user. Then when the day was finished, I wrote the last two impressions for every user to validation and everything else to the remainder of the training set.\r\n\r\nDoes anyone spot an obvious flaw in this strategy?\r\nHave you been able to use any user features directly without suffering a similar issue?  What did you change to the model and learning to be able to use them?\r\n\r\nEDIT: changed b18/19 -> c21/22",
      "votes": null
    },
    {
      "id": "87418",
      "postDate": "07/29/2015 12:58:34",
      "content": "<p>@remap,</p>\n\n<p>Let me take a stab at what your model may have been missing.  If you look carefully at the test data, around 27% of the searches come from new users.  Additionally, I think around 10-15% of searches in the training data come from users that only have one search in training.  So any model features based on historical user data won't work without a proper validation set.  Additionally, around 10-15% of the users make up 40% of the searches in the training data.  So here's the imbalance.  If test only has one search per user, but train has many, then unless your validation set mimics the one search per user and has roughly 27% of searches on users you haven't trained on, it is not going to give you a good signal.</p>\n\n<p>As for only training on clean global time (Don't train on 5/20 for a 5/19 prediction), we only held this per user, but between users didn't stick to this.  I know this setup is a bit unrealistic, however training on 5/19 &amp; 5/20 for different users that are predicted for 5/18 works in this competition because every users' time schedule is different.</p>\n\n<p>In the end, userid was an important feature for us.  It is in namespace &quot;F&quot; in the parameters I showed at: <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations/87373#post87373\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations/87373#post87373</a></p>",
      "rawMarkdown": "remap,\r\n\r\nLet me take a stab at what your model may have been missing.  If you look carefully at the test data, around 27% of the searches come from new users.  Additionally, I think around 10-15% of searches in the training data come from users that only have one search in training.  So any model features based on historical user data won't work without a proper validation set.  Additionally, around 10-15% of the users make up 40% of the searches in the training data.  So here's the imbalance.  If test only has one search per user, but train has many, then unless your validation set mimics the one search per user and has roughly 27% of searches on users you haven't trained on, it is not going to give you a good signal.\r\n\r\nAs for only training on clean global time (Don't train on 5/20 for a 5/19 prediction), we only held this per user, but between users didn't stick to this.  I know this setup is a bit unrealistic, however training on 5/19 & 5/20 for different users that are predicted for 5/18 works in this competition because every users' time schedule is different.\r\n\r\nIn the end, userid was an important feature for us.  It is in namespace \"F\" in the parameters I showed at: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations/87373#post87373",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 87418,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "07/29/2015 12:58:34",
      "content": "<p>@remap,</p>\n\n<p>Let me take a stab at what your model may have been missing.  If you look carefully at the test data, around 27% of the searches come from new users.  Additionally, I think around 10-15% of searches in the training data come from users that only have one search in training.  So any model features based on historical user data won't work without a proper validation set.  Additionally, around 10-15% of the users make up 40% of the searches in the training data.  So here's the imbalance.  If test only has one search per user, but train has many, then unless your validation set mimics the one search per user and has roughly 27% of searches on users you haven't trained on, it is not going to give you a good signal.</p>\n\n<p>As for only training on clean global time (Don't train on 5/20 for a 5/19 prediction), we only held this per user, but between users didn't stick to this.  I know this setup is a bit unrealistic, however training on 5/19 &amp; 5/20 for different users that are predicted for 5/18 works in this competition because every users' time schedule is different.</p>\n\n<p>In the end, userid was an important feature for us.  It is in namespace &quot;F&quot; in the parameters I showed at: <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations/87373#post87373\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations/87373#post87373</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "87407": "Congrats to those who won. \r\n\r\nSo I was cursed in this competition for some reason; my validation score blew up when I used anything related to a user. This is no longer relevant for the competition, but I'd really like to know what I did wrong for such a basic issue.\r\n\r\nI had training data lines like this:\r\n\r\n    -1 1 |u u1:570979 u2:4159 u3:63 u4:3 u5:3566 u6:0 u7:53 |a a8:16904117 a9:7 a10:38 a11:3 a12:41511.0 |s s13:131 s14:190 s15:-1 s16:-1 |b b18:2 b19:26 |c c21:26 c22:18 c23:1 |h h24:0.002913\r\n\r\n|u are things from the userinfo table, |a adsinfo data, |s day/hour/city/region, |b ad statistics, |c user statistics, |h histctr.\r\n\r\nThe user statistics are not directly statistically related; c21 is a representation of the click through rate adjusted for users that have really small numbers of clicks using an algorithm similar to a wilson score. Low number of clicks keep the user as \"average\" and anything >10 already distinguishes this user from others. c22 is simply the number of days the user prowls the site. c23 is only used to figure out if a user appears in the trainstream or not.\r\n\r\nI could selectively disable any feature. Without |c features (user), I got cv results that approximated what I got on the LB, but huge differences when any user feature got activated. I only used FTRL. Example:\r\n\r\nwithout user features c21 and c22:\r\ntraining: 0.035, cv: 0.046, LB: 0.045\r\nwith user features c21 and c22:\r\ntraining: 0.029, cv: 0.038, LB: 0.061\r\n\r\nTraining data was sorted by date, then hour. The exact same code for training, validation and test got used each time.\r\n\r\nFrom the entire training set, I collected impressions for every user and stored them separately for every user. Then when the day was finished, I wrote the last two impressions for every user to validation and everything else to the remainder of the training set.\r\n\r\nDoes anyone spot an obvious flaw in this strategy?\r\nHave you been able to use any user features directly without suffering a similar issue?  What did you change to the model and learning to be able to use them?\r\n\r\nEDIT: changed b18/19 -> c21/22",
    "87418": "remap,\r\n\r\nLet me take a stab at what your model may have been missing.  If you look carefully at the test data, around 27% of the searches come from new users.  Additionally, I think around 10-15% of searches in the training data come from users that only have one search in training.  So any model features based on historical user data won't work without a proper validation set.  Additionally, around 10-15% of the users make up 40% of the searches in the training data.  So here's the imbalance.  If test only has one search per user, but train has many, then unless your validation set mimics the one search per user and has roughly 27% of searches on users you haven't trained on, it is not going to give you a good signal.\r\n\r\nAs for only training on clean global time (Don't train on 5/20 for a 5/19 prediction), we only held this per user, but between users didn't stick to this.  I know this setup is a bit unrealistic, however training on 5/19 & 5/20 for different users that are predicted for 5/18 works in this competition because every users' time schedule is different.\r\n\r\nIn the end, userid was an important feature for us.  It is in namespace \"F\" in the parameters I showed at: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15606/congratulations/87373#post87373"
  },
  "source": "meta"
}