{
  "id": 15367,
  "title": "Proper validation set",
  "url": "/competitions/avito-context-ad-clicks/discussion/15367",
  "author_name": "",
  "post_date": "2015-07-19T14:40:30.863Z",
  "votes": 14,
  "comment_count": 6,
  "views": 1702,
  "content": "<p>Guys, I see many people struggling to get a good validation set so I decided to help:</p>\n\n<ol>\n<li>Take the minimum date that occurs in the test set (May 12?).</li>\n<li><p>Take all LAST sessions for each user that happens after that date (or sample it). The last session is the one with highest date for each user.</p>\n\n<p>This problem has a time component, so sampling from other periods of time won't do, since it is expected that ads change over time.\nRandomly sampling wont do, since it will get a bigger proportion of users that has many sessions. \nIf we take anything but the last (or n last) sessions from user we will leave the user future in the training set, so cv wont match because of overfit.\nBottom line: In the data page it says that they take the last session for each user that happens after may 12. So we just need to do the same (as close as it gets)! </p></li>\n</ol>\n\n<p>Happy validation!</p>",
  "messages": [
    {
      "id": "86077",
      "postDate": "07/19/2015 14:40:30",
      "content": "<p>Guys, I see many people struggling to get a good validation set so I decided to help:</p>\n\n<ol>\n<li>Take the minimum date that occurs in the test set (May 12?).</li>\n<li><p>Take all LAST sessions for each user that happens after that date (or sample it). The last session is the one with highest date for each user.</p>\n\n<p>This problem has a time component, so sampling from other periods of time won't do, since it is expected that ads change over time.\nRandomly sampling wont do, since it will get a bigger proportion of users that has many sessions. \nIf we take anything but the last (or n last) sessions from user we will leave the user future in the training set, so cv wont match because of overfit.\nBottom line: In the data page it says that they take the last session for each user that happens after may 12. So we just need to do the same (as close as it gets)! </p></li>\n</ol>\n\n<p>Happy validation!</p>",
      "rawMarkdown": "Guys, I see many people struggling to get a good validation set so I decided to help:\r\n\r\n 1. Take the minimum date that occurs in the test set (May 12?).\r\n 2. Take all LAST sessions for each user that happens after that date (or sample it). The last session is the one with highest date for each user.\r\n\r\n    This problem has a time component, so sampling from other periods of time won't do, since it is expected that ads change over time.\r\n    Randomly sampling wont do, since it will get a bigger proportion of users that has many sessions. \r\n    If we take anything but the last (or n last) sessions from user we will leave the user future in the training set, so cv wont match because of overfit.\r\n    Bottom line: In the data page it says that they take the last session for each user that happens after may 12. So we just need to do the same (as close as it gets)! \r\n\r\nHappy validation!",
      "votes": null
    },
    {
      "id": "86079",
      "postDate": "07/19/2015 14:56:28",
      "content": "<p>Thank you! I figured out this several days ago based on Leustagos' another post : ) \nI just want to confirm that it works very well.</p>",
      "rawMarkdown": "Thank you! I figured out this several days ago based on Leustagos' another post : ) \r\nI just want to confirm that it works very well.",
      "votes": null
    },
    {
      "id": "86207",
      "postDate": "07/20/2015 22:24:03",
      "content": "<p>That's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. </p>\n\n<p>When I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.</p>\n\n<p>Meaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).</p>",
      "rawMarkdown": "That's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. \r\n\r\nWhen I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.\r\n\r\nMeaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).",
      "votes": null
    },
    {
      "id": "86209",
      "postDate": "07/20/2015 22:32:33",
      "content": "<p>[quote=Remap on github;86207]</p>\n\n<p>That's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. </p>\n\n<p>When I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.</p>\n\n<p>Meaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).</p>\n\n<p>[/quote]\n You are overfitting. The moment you started using the outputs of the validation set to build features, it stopped being a proper validation set. Thats all. Never use the output of the validation set if you want it to stay accurate (never is quite extreme as there ways to mitigate it, but is too complex to discuss here).\nRecalculate your user features excluding the validation set and you will see that the validation score will skyrocket too.</p>",
      "rawMarkdown": "[quote=Remap on github;86207]\r\n\r\nThat's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. \r\n\r\nWhen I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.\r\n\r\nMeaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).\r\n\r\n[/quote]\r\n You are overfitting. The moment you started using the outputs of the validation set to build features, it stopped being a proper validation set. Thats all. Never use the output of the validation set if you want it to stay accurate (never is quite extreme as there ways to mitigate it, but is too complex to discuss here).\r\nRecalculate your user features excluding the validation set and you will see that the validation score will skyrocket too.",
      "votes": null
    },
    {
      "id": "86211",
      "postDate": "07/20/2015 23:00:17",
      "content": "<p>Ok, so by my knowledge I'm not using any output from the validation set. </p>\n\n<p>My training set is built on every data record &lt; day 132, then I distribute records after 132 between training and validation by a random choice algorithm. I do make sure that the validation set has at least one record per user after 132.</p>\n\n<p>I calculate statistics the following way:\n- Used the training set (not validation!) to calculate user statistics, which are used to train, validate then test. train = 0.029, validate= 0.038, test was 0.061\n- Used the entire dataset to calculate user stats, which are used as features to train, validate and test. Train was 0.028, validate was 0.037, test was 0.065. (worse).</p>\n\n<p>In my implementation, I can selectively disable features. With the full set enabled, I get results above. When I selectively disable &quot;user&quot; features (whilst keeping ad features and averages enabled), I then get results where the LB results get pulled down to validation score level.</p>\n\n<p>From what you're saying is that I should recalculate my user features excluding records &gt; 132?</p>",
      "rawMarkdown": "Ok, so by my knowledge I'm not using any output from the validation set. \r\n\r\nMy training set is built on every data record < day 132, then I distribute records after 132 between training and validation by a random choice algorithm. I do make sure that the validation set has at least one record per user after 132.\r\n\r\nI calculate statistics the following way:\r\n- Used the training set (not validation!) to calculate user statistics, which are used to train, validate then test. train = 0.029, validate= 0.038, test was 0.061\r\n- Used the entire dataset to calculate user stats, which are used as features to train, validate and test. Train was 0.028, validate was 0.037, test was 0.065. (worse).\r\n\r\nIn my implementation, I can selectively disable features. With the full set enabled, I get results above. When I selectively disable \"user\" features (whilst keeping ad features and averages enabled), I then get results where the LB results get pulled down to validation score level.\r\n\r\nFrom what you're saying is that I should recalculate my user features excluding records > 132?",
      "votes": null
    },
    {
      "id": "86213",
      "postDate": "07/20/2015 23:04:15",
      "content": "<p>From the results you are reporting, you must be using outputs from validation or making features that doesnt happen the same in the leaderboard.</p>",
      "rawMarkdown": "From the results you are reporting, you must be using outputs from validation or making features that doesnt happen the same in the leaderboard.",
      "votes": null
    },
    {
      "id": "86341",
      "postDate": "07/21/2015 17:38:14",
      "content": "<p>There's 10% of users that do not appear in search streams whatsoever (not in my training anyway), so those are users that have landed on the site through google or whatever and only interacted through visits/phone interactions. I'm also processing data sequentially by date and there are users that appear late in the dataset. When learning has saturated for some features, those users probably suffer a very large penalty if their features are individualistic, not categorical. Reducing learning rate doesn't help in those cases, because of the way how they appear in the dataset and their relative penalty doesn't change by setting the learning rate. The best way I think is to re-engineer the features (either that, or there are hash collisions, but I'm using a really good hash). </p>",
      "rawMarkdown": "There's 10% of users that do not appear in search streams whatsoever (not in my training anyway), so those are users that have landed on the site through google or whatever and only interacted through visits/phone interactions. I'm also processing data sequentially by date and there are users that appear late in the dataset. When learning has saturated for some features, those users probably suffer a very large penalty if their features are individualistic, not categorical. Reducing learning rate doesn't help in those cases, because of the way how they appear in the dataset and their relative penalty doesn't change by setting the learning rate. The best way I think is to re-engineer the features (either that, or there are hash collisions, but I'm using a really good hash).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 86079,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "07/19/2015 14:56:28",
      "content": "<p>Thank you! I figured out this several days ago based on Leustagos' another post : ) \nI just want to confirm that it works very well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86207,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/20/2015 22:24:03",
      "content": "<p>That's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. </p>\n\n<p>When I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.</p>\n\n<p>Meaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86209,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/20/2015 22:32:33",
      "content": "<p>[quote=Remap on github;86207]</p>\n\n<p>That's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. </p>\n\n<p>When I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.</p>\n\n<p>Meaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).</p>\n\n<p>[/quote]\n You are overfitting. The moment you started using the outputs of the validation set to build features, it stopped being a proper validation set. Thats all. Never use the output of the validation set if you want it to stay accurate (never is quite extreme as there ways to mitigate it, but is too complex to discuss here).\nRecalculate your user features excluding the validation set and you will see that the validation score will skyrocket too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86211,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/20/2015 23:00:17",
      "content": "<p>Ok, so by my knowledge I'm not using any output from the validation set. </p>\n\n<p>My training set is built on every data record &lt; day 132, then I distribute records after 132 between training and validation by a random choice algorithm. I do make sure that the validation set has at least one record per user after 132.</p>\n\n<p>I calculate statistics the following way:\n- Used the training set (not validation!) to calculate user statistics, which are used to train, validate then test. train = 0.029, validate= 0.038, test was 0.061\n- Used the entire dataset to calculate user stats, which are used as features to train, validate and test. Train was 0.028, validate was 0.037, test was 0.065. (worse).</p>\n\n<p>In my implementation, I can selectively disable features. With the full set enabled, I get results above. When I selectively disable &quot;user&quot; features (whilst keeping ad features and averages enabled), I then get results where the LB results get pulled down to validation score level.</p>\n\n<p>From what you're saying is that I should recalculate my user features excluding records &gt; 132?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86213,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/20/2015 23:04:15",
      "content": "<p>From the results you are reporting, you must be using outputs from validation or making features that doesnt happen the same in the leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86341,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/21/2015 17:38:14",
      "content": "<p>There's 10% of users that do not appear in search streams whatsoever (not in my training anyway), so those are users that have landed on the site through google or whatever and only interacted through visits/phone interactions. I'm also processing data sequentially by date and there are users that appear late in the dataset. When learning has saturated for some features, those users probably suffer a very large penalty if their features are individualistic, not categorical. Reducing learning rate doesn't help in those cases, because of the way how they appear in the dataset and their relative penalty doesn't change by setting the learning rate. The best way I think is to re-engineer the features (either that, or there are hash collisions, but I'm using a really good hash). </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "86077": "Guys, I see many people struggling to get a good validation set so I decided to help:\r\n\r\n 1. Take the minimum date that occurs in the test set (May 12?).\r\n 2. Take all LAST sessions for each user that happens after that date (or sample it). The last session is the one with highest date for each user.\r\n\r\n    This problem has a time component, so sampling from other periods of time won't do, since it is expected that ads change over time.\r\n    Randomly sampling wont do, since it will get a bigger proportion of users that has many sessions. \r\n    If we take anything but the last (or n last) sessions from user we will leave the user future in the training set, so cv wont match because of overfit.\r\n    Bottom line: In the data page it says that they take the last session for each user that happens after may 12. So we just need to do the same (as close as it gets)! \r\n\r\nHappy validation!",
    "86079": "Thank you! I figured out this several days ago based on Leustagos' another post : ) \r\nI just want to confirm that it works very well.",
    "86207": "That's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. \r\n\r\nWhen I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.\r\n\r\nMeaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).",
    "86209": "[quote=Remap on github;86207]\r\n\r\nThat's not all of it though :). When I use user features, the validation set reports 0.038 and my training set 0.029, yet my LB submissions go straight up to 0.06+ or so. \r\n\r\nWhen I cancel out the user features, the correlation between submissions and local validation returns and I get very small deviations. The only thing I did for users are calculating average impressions and average clicks over the days that these users are active. Apparently there are too many new users to the site or they started behaving completely differently.\r\n\r\nMeaning... it's not only the validation set that matters, as I learned over the past 15 days :), it's certainly also the model you use (the features that make it).\r\n\r\n[/quote]\r\n You are overfitting. The moment you started using the outputs of the validation set to build features, it stopped being a proper validation set. Thats all. Never use the output of the validation set if you want it to stay accurate (never is quite extreme as there ways to mitigate it, but is too complex to discuss here).\r\nRecalculate your user features excluding the validation set and you will see that the validation score will skyrocket too.",
    "86211": "Ok, so by my knowledge I'm not using any output from the validation set. \r\n\r\nMy training set is built on every data record < day 132, then I distribute records after 132 between training and validation by a random choice algorithm. I do make sure that the validation set has at least one record per user after 132.\r\n\r\nI calculate statistics the following way:\r\n- Used the training set (not validation!) to calculate user statistics, which are used to train, validate then test. train = 0.029, validate= 0.038, test was 0.061\r\n- Used the entire dataset to calculate user stats, which are used as features to train, validate and test. Train was 0.028, validate was 0.037, test was 0.065. (worse).\r\n\r\nIn my implementation, I can selectively disable features. With the full set enabled, I get results above. When I selectively disable \"user\" features (whilst keeping ad features and averages enabled), I then get results where the LB results get pulled down to validation score level.\r\n\r\nFrom what you're saying is that I should recalculate my user features excluding records > 132?",
    "86213": "From the results you are reporting, you must be using outputs from validation or making features that doesnt happen the same in the leaderboard.",
    "86341": "There's 10% of users that do not appear in search streams whatsoever (not in my training anyway), so those are users that have landed on the site through google or whatever and only interacted through visits/phone interactions. I'm also processing data sequentially by date and there are users that appear late in the dataset. When learning has saturated for some features, those users probably suffer a very large penalty if their features are individualistic, not categorical. Reducing learning rate doesn't help in those cases, because of the way how they appear in the dataset and their relative penalty doesn't change by setting the learning rate. The best way I think is to re-engineer the features (either that, or there are hash collisions, but I'm using a really good hash)."
  },
  "source": "meta"
}