{
  "id": 15089,
  "title": "Cross validation error higher against leaderboard",
  "url": "/competitions/avito-context-ad-clicks/discussion/15089",
  "author_name": "",
  "post_date": "2015-07-07T12:35:43.067Z",
  "votes": null,
  "comment_count": 9,
  "views": 1486,
  "content": "<p>Just getting started with machine learning. I did the netflix prize way back and want to get back into machine learning as a hobby.</p>\n\n<p>So my training set I'm using is roughly 190M records. I take a random sample out of that of anything between 1-10M records. With sklearn and a 10-fold cross validation set, each iteration trains the model from scratch.</p>\n\n<p>Running each trained model against each test fold, I get really good numbers of 0.036 for all 10 test folds, +/- 0.002 (0.034-0.038). </p>\n\n<p>Then I train the entire model against the entire training set, do the predictions, yet the calculated score on the leaderboard is consistently 0.01 higher (0.047-ish). All records are there and there are no seemingly invalid predictions (outside [0,1]).</p>\n\n<p>This is my logloss function to calculate the results:</p>\n\n<pre><code>def llfun(act, pred):\n    epsilon = 1e-15\n    pred = sp.maximum(epsilon, pred)\n    pred = sp.minimum(1-epsilon, pred)\n    ll = sum(act*sp.log(pred) + sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\n    ll = ll * -1.0/len(act)\n    return ll\n</code></pre>\n\n<p>So my cross validation results aren't representative for the test data, even though none of the test data was used in training and all test folds have roughly similar total errors.</p>\n\n<ul>\n<li>Do you get representative total errors (vs. leaderboard calculated result) in your cross validation?</li>\n<li>Could this be due to the test date span != train date span?</li>\n</ul>",
  "messages": [
    {
      "id": "83653",
      "postDate": "07/07/2015 12:35:43",
      "content": "<p>Just getting started with machine learning. I did the netflix prize way back and want to get back into machine learning as a hobby.</p>\n\n<p>So my training set I'm using is roughly 190M records. I take a random sample out of that of anything between 1-10M records. With sklearn and a 10-fold cross validation set, each iteration trains the model from scratch.</p>\n\n<p>Running each trained model against each test fold, I get really good numbers of 0.036 for all 10 test folds, +/- 0.002 (0.034-0.038). </p>\n\n<p>Then I train the entire model against the entire training set, do the predictions, yet the calculated score on the leaderboard is consistently 0.01 higher (0.047-ish). All records are there and there are no seemingly invalid predictions (outside [0,1]).</p>\n\n<p>This is my logloss function to calculate the results:</p>\n\n<pre><code>def llfun(act, pred):\n    epsilon = 1e-15\n    pred = sp.maximum(epsilon, pred)\n    pred = sp.minimum(1-epsilon, pred)\n    ll = sum(act*sp.log(pred) + sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\n    ll = ll * -1.0/len(act)\n    return ll\n</code></pre>\n\n<p>So my cross validation results aren't representative for the test data, even though none of the test data was used in training and all test folds have roughly similar total errors.</p>\n\n<ul>\n<li>Do you get representative total errors (vs. leaderboard calculated result) in your cross validation?</li>\n<li>Could this be due to the test date span != train date span?</li>\n</ul>",
      "rawMarkdown": "Just getting started with machine learning. I did the netflix prize way back and want to get back into machine learning as a hobby.\r\n\r\nSo my training set I'm using is roughly 190M records. I take a random sample out of that of anything between 1-10M records. With sklearn and a 10-fold cross validation set, each iteration trains the model from scratch.\r\n\r\nRunning each trained model against each test fold, I get really good numbers of 0.036 for all 10 test folds, +/- 0.002 (0.034-0.038). \r\n\r\nThen I train the entire model against the entire training set, do the predictions, yet the calculated score on the leaderboard is consistently 0.01 higher (0.047-ish). All records are there and there are no seemingly invalid predictions (outside [0,1]).\r\n\r\nThis is my logloss function to calculate the results:\r\n\r\n    def llfun(act, pred):\r\n        epsilon = 1e-15\r\n        pred = sp.maximum(epsilon, pred)\r\n        pred = sp.minimum(1-epsilon, pred)\r\n        ll = sum(act*sp.log(pred) + sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\r\n        ll = ll * -1.0/len(act)\r\n        return ll\r\n\r\nSo my cross validation results aren't representative for the test data, even though none of the test data was used in training and all test folds have roughly similar total errors.\r\n\r\n - Do you get representative total errors (vs. leaderboard calculated result) in your cross validation?\r\n - Could this be due to the test date span != train date span?",
      "votes": null
    },
    {
      "id": "83665",
      "postDate": "07/07/2015 14:38:32",
      "content": "<p>Your model and code seems to be ok!\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...</p>",
      "rawMarkdown": "Your model and code seems to be ok!\r\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...",
      "votes": null
    },
    {
      "id": "83669",
      "postDate": "07/07/2015 15:07:47",
      "content": "<p>Ah: <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14580/histctr\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14580/histctr</a></p>\n\n<p>As explained by Gert. Which probably means that the site has reacted to user's search patterns and presented more relevant results based on the behavior in the training set.</p>",
      "rawMarkdown": "Ah: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14580/histctr\r\n\r\nAs explained by Gert. Which probably means that the site has reacted to user's search patterns and presented more relevant results based on the behavior in the training set.",
      "votes": null
    },
    {
      "id": "83682",
      "postDate": "07/07/2015 18:10:19",
      "content": "<p>[quote=Gilberto Titericz Junior;83665]</p>\n\n<p>Your model and code seems to be ok!\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...</p>\n\n<p>[/quote]</p>\n\n<p>Gilberto, is this CV with k-fold (k &gt;1) or based in a time splitting with only a held data set?</p>\n\n<p>By the way, can someone say me where I can fork Owen's script :-)</p>",
      "rawMarkdown": "[quote=Gilberto Titericz Junior;83665]\r\n\r\nYour model and code seems to be ok!\r\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...\r\n\r\n\r\n\r\n[/quote]\r\n\r\nGilberto, is this CV with k-fold (k >1) or based in a time splitting with only a held data set?\r\n\r\nBy the way, can someone say me where I can fork Owen's script :-)",
      "votes": null
    },
    {
      "id": "83703",
      "postDate": "07/07/2015 21:16:32",
      "content": "<p>[quote=Jos&#233; A. Guerrero;83682]</p>\n\n<p>[quote=Gilberto Titericz Junior;83665]</p>\n\n<p>Your model and code seems to be ok!\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...</p>\n\n<p>[/quote]</p>\n\n<p>Gilberto, is this CV with k-fold (k &gt;1) or based in a time splitting with only a held data set?</p>\n\n<p>By the way, can someone say me where I can fork Owen's script :-)</p>\n\n<p>[/quote]</p>\n\n<p>Jos&#233;: start here, it may help...\n<a href=\"https://github.com/owenzhang/kaggle-avazu\">https://github.com/owenzhang/kaggle-avazu</a></p>\n\n<p>My cv is consistent with LB. Usually ~0.001 higher. And cv improvments reflects lb improvments.</p>\n\n<p>For 0.04077 on lb, my cv is 0.04250</p>",
      "rawMarkdown": "[quote=José A. Guerrero;83682]\r\n\r\n[quote=Gilberto Titericz Junior;83665]\r\n\r\nYour model and code seems to be ok!\r\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...\r\n\r\n\r\n\r\n[/quote]\r\n\r\nGilberto, is this CV with k-fold (k >1) or based in a time splitting with only a held data set?\r\n\r\nBy the way, can someone say me where I can fork Owen's script :-)\r\n\r\n\r\n\r\n\r\n\r\n[/quote]\r\n\r\n\r\nJosé: start here, it may help...\r\nhttps://github.com/owenzhang/kaggle-avazu\r\n\r\nMy cv is consistent with LB. Usually ~0.001 higher. And cv improvments reflects lb improvments.\r\n\r\nFor 0.04077 on lb, my cv is 0.04250",
      "votes": null
    },
    {
      "id": "83908",
      "postDate": "07/09/2015 08:45:33",
      "content": "<p>The possible cause of the difference is a way the dataset was generated. As it is describes in &quot;Data Splitting&quot; section, the data is a result of the following steps:</p>\n\n<ol>\n<li>select test time interval (May 12 - May 20)</li>\n<li>take random sample of users having activity between May 12 and May 20</li>\n<li>for each user select a test point (target impression), make the test set from these points</li>\n<li>record all activities of the selected users in previous month and make the train set from it.</li>\n</ol>\n\n<p>The problem is that now train set contains only those users that will be active at test time. The train set misses the users who get on the site, make one-three searches and leave it forever. And it seems that the latter ones behave differently than the 'regular' users with long search histories.</p>\n\n<p>If you want to obtain a train set similar to test one, make it from the last user impressions before the test points. The HistCRT score on such set is 0.05747 which is quite closed to the 0.05717 of test set.</p>",
      "rawMarkdown": "The possible cause of the difference is a way the dataset was generated. As it is describes in \"Data Splitting\" section, the data is a result of the following steps:\r\n\r\n 1. select test time interval (May 12 - May 20)\r\n 2. take random sample of users having activity between May 12 and May 20\r\n 3. for each user select a test point (target impression), make the test set from these points\r\n 4. record all activities of the selected users in previous month and make the train set from it.\r\n\r\nThe problem is that now train set contains only those users that will be active at test time. The train set misses the users who get on the site, make one-three searches and leave it forever. And it seems that the latter ones behave differently than the 'regular' users with long search histories.\r\n\r\nIf you want to obtain a train set similar to test one, make it from the last user impressions before the test points. The HistCRT score on such set is 0.05747 which is quite closed to the 0.05717 of test set.",
      "votes": null
    },
    {
      "id": "84096",
      "postDate": "07/11/2015 02:10:38",
      "content": "<p>I thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. </p>\n\n<p>I'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.</p>\n\n<p>What you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.</p>\n\n<p>So my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.</p>\n\n<p>Other people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.</p>",
      "rawMarkdown": "I thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. \r\n\r\nI'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.\r\n\r\nWhat you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.\r\n\r\nSo my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.\r\n\r\nOther people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.",
      "votes": null
    },
    {
      "id": "84103",
      "postDate": "07/11/2015 04:33:43",
      "content": "<p>By reading the Data Split section (last one) on the data page (<a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/data\">https://www.kaggle.com/c/avito-context-ad-clicks/data</a>). Its all there. Just make validation and training have the exact same relation that lb and the full train have. If I say more i'll just be handing the solution (actually I think I just did by saying so much)</p>\n\n<p>[quote=Remap on github;84096]</p>\n\n<p>I thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. </p>\n\n<p>I'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.</p>\n\n<p>What you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.</p>\n\n<p>So my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.</p>\n\n<p>Other people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "By reading the Data Split section (last one) on the data page (https://www.kaggle.com/c/avito-context-ad-clicks/data). Its all there. Just make validation and training have the exact same relation that lb and the full train have. If I say more i'll just be handing the solution (actually I think I just did by saying so much)\r\n\r\n[quote=Remap on github;84096]\r\n\r\nI thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. \r\n\r\nI'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.\r\n\r\nWhat you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.\r\n\r\nSo my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.\r\n\r\nOther people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "84132",
      "postDate": "07/11/2015 12:39:51",
      "content": "<p>Ok, so I process everything with flat files, because they're a lot faster to remodel data than SQL queries. It is indeed extremely important to pay very close attention to every detail here. I was trying to cut this short by simply tailing the end of the training file, but that doesn't create the correct distribution.</p>\n\n<p>Previously the validation scored 0.041 on the very first run of 1M records. I think now I have the correct data to validate against. It doesn't change my position, but at least now I can play around more and get a feel for what changes in parameters really do.</p>\n\n<p>Thank you very much!</p>\n\n<pre><code>Validation stage\n( (1000000, 0.043459011918673934, 0.043459011918673934))\n( (2000000, 0.04460827962764392, 0.0457575473366139))\n( (3000000, 0.04665980910677317, 0.05076286806503169))\n( (4000000, 0.04773306449442931, 0.05095283065739771))\n( (5000000, 0.04889447220230085, 0.053540103033787007))\n( (5459321, 0.04964257027849093))\n</code></pre>",
      "rawMarkdown": "Ok, so I process everything with flat files, because they're a lot faster to remodel data than SQL queries. It is indeed extremely important to pay very close attention to every detail here. I was trying to cut this short by simply tailing the end of the training file, but that doesn't create the correct distribution.\r\n\r\nPreviously the validation scored 0.041 on the very first run of 1M records. I think now I have the correct data to validate against. It doesn't change my position, but at least now I can play around more and get a feel for what changes in parameters really do.\r\n\r\nThank you very much!\r\n\r\n    Validation stage\r\n    ( (1000000, 0.043459011918673934, 0.043459011918673934))\r\n    ( (2000000, 0.04460827962764392, 0.0457575473366139))\r\n    ( (3000000, 0.04665980910677317, 0.05076286806503169))\r\n    ( (4000000, 0.04773306449442931, 0.05095283065739771))\r\n    ( (5000000, 0.04889447220230085, 0.053540103033787007))\r\n    ( (5459321, 0.04964257027849093))",
      "votes": null
    },
    {
      "id": "84157",
      "postDate": "07/11/2015 19:12:20",
      "content": "<p>Looks great!</p>\n\n<p>A test run on 5% of the data:  0.047371</p>\n\n<p>LB submission: 0.04734</p>",
      "rawMarkdown": "Looks great!\r\n\r\nA test run on 5% of the data:  0.047371\r\n\r\nLB submission: 0.04734",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 83665,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/07/2015 14:38:32",
      "content": "<p>Your model and code seems to be ok!\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83669,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/07/2015 15:07:47",
      "content": "<p>Ah: <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14580/histctr\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14580/histctr</a></p>\n\n<p>As explained by Gert. Which probably means that the site has reacted to user's search patterns and presented more relevant results based on the behavior in the training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83682,
      "author_name": "blindape",
      "author_url": "",
      "post_date": "07/07/2015 18:10:19",
      "content": "<p>[quote=Gilberto Titericz Junior;83665]</p>\n\n<p>Your model and code seems to be ok!\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...</p>\n\n<p>[/quote]</p>\n\n<p>Gilberto, is this CV with k-fold (k &gt;1) or based in a time splitting with only a held data set?</p>\n\n<p>By the way, can someone say me where I can fork Owen's script :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83703,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/07/2015 21:16:32",
      "content": "<p>[quote=Jos&#233; A. Guerrero;83682]</p>\n\n<p>[quote=Gilberto Titericz Junior;83665]</p>\n\n<p>Your model and code seems to be ok!\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...</p>\n\n<p>[/quote]</p>\n\n<p>Gilberto, is this CV with k-fold (k &gt;1) or based in a time splitting with only a held data set?</p>\n\n<p>By the way, can someone say me where I can fork Owen's script :-)</p>\n\n<p>[/quote]</p>\n\n<p>Jos&#233;: start here, it may help...\n<a href=\"https://github.com/owenzhang/kaggle-avazu\">https://github.com/owenzhang/kaggle-avazu</a></p>\n\n<p>My cv is consistent with LB. Usually ~0.001 higher. And cv improvments reflects lb improvments.</p>\n\n<p>For 0.04077 on lb, my cv is 0.04250</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83908,
      "author_name": "skirpichenko",
      "author_url": "",
      "post_date": "07/09/2015 08:45:33",
      "content": "<p>The possible cause of the difference is a way the dataset was generated. As it is describes in &quot;Data Splitting&quot; section, the data is a result of the following steps:</p>\n\n<ol>\n<li>select test time interval (May 12 - May 20)</li>\n<li>take random sample of users having activity between May 12 and May 20</li>\n<li>for each user select a test point (target impression), make the test set from these points</li>\n<li>record all activities of the selected users in previous month and make the train set from it.</li>\n</ol>\n\n<p>The problem is that now train set contains only those users that will be active at test time. The train set misses the users who get on the site, make one-three searches and leave it forever. And it seems that the latter ones behave differently than the 'regular' users with long search histories.</p>\n\n<p>If you want to obtain a train set similar to test one, make it from the last user impressions before the test points. The HistCRT score on such set is 0.05747 which is quite closed to the 0.05717 of test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84096,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/11/2015 02:10:38",
      "content": "<p>I thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. </p>\n\n<p>I'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.</p>\n\n<p>What you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.</p>\n\n<p>So my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.</p>\n\n<p>Other people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84103,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/11/2015 04:33:43",
      "content": "<p>By reading the Data Split section (last one) on the data page (<a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/data\">https://www.kaggle.com/c/avito-context-ad-clicks/data</a>). Its all there. Just make validation and training have the exact same relation that lb and the full train have. If I say more i'll just be handing the solution (actually I think I just did by saying so much)</p>\n\n<p>[quote=Remap on github;84096]</p>\n\n<p>I thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. </p>\n\n<p>I'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.</p>\n\n<p>What you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.</p>\n\n<p>So my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.</p>\n\n<p>Other people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84132,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/11/2015 12:39:51",
      "content": "<p>Ok, so I process everything with flat files, because they're a lot faster to remodel data than SQL queries. It is indeed extremely important to pay very close attention to every detail here. I was trying to cut this short by simply tailing the end of the training file, but that doesn't create the correct distribution.</p>\n\n<p>Previously the validation scored 0.041 on the very first run of 1M records. I think now I have the correct data to validate against. It doesn't change my position, but at least now I can play around more and get a feel for what changes in parameters really do.</p>\n\n<p>Thank you very much!</p>\n\n<pre><code>Validation stage\n( (1000000, 0.043459011918673934, 0.043459011918673934))\n( (2000000, 0.04460827962764392, 0.0457575473366139))\n( (3000000, 0.04665980910677317, 0.05076286806503169))\n( (4000000, 0.04773306449442931, 0.05095283065739771))\n( (5000000, 0.04889447220230085, 0.053540103033787007))\n( (5459321, 0.04964257027849093))\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84157,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/11/2015 19:12:20",
      "content": "<p>Looks great!</p>\n\n<p>A test run on 5% of the data:  0.047371</p>\n\n<p>LB submission: 0.04734</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "83653": "Just getting started with machine learning. I did the netflix prize way back and want to get back into machine learning as a hobby.\r\n\r\nSo my training set I'm using is roughly 190M records. I take a random sample out of that of anything between 1-10M records. With sklearn and a 10-fold cross validation set, each iteration trains the model from scratch.\r\n\r\nRunning each trained model against each test fold, I get really good numbers of 0.036 for all 10 test folds, +/- 0.002 (0.034-0.038). \r\n\r\nThen I train the entire model against the entire training set, do the predictions, yet the calculated score on the leaderboard is consistently 0.01 higher (0.047-ish). All records are there and there are no seemingly invalid predictions (outside [0,1]).\r\n\r\nThis is my logloss function to calculate the results:\r\n\r\n    def llfun(act, pred):\r\n        epsilon = 1e-15\r\n        pred = sp.maximum(epsilon, pred)\r\n        pred = sp.minimum(1-epsilon, pred)\r\n        ll = sum(act*sp.log(pred) + sp.subtract(1,act)*sp.log(sp.subtract(1,pred)))\r\n        ll = ll * -1.0/len(act)\r\n        return ll\r\n\r\nSo my cross validation results aren't representative for the test data, even though none of the test data was used in training and all test folds have roughly similar total errors.\r\n\r\n - Do you get representative total errors (vs. leaderboard calculated result) in your cross validation?\r\n - Could this be due to the test date span != train date span?",
    "83665": "Your model and code seems to be ok!\r\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...",
    "83669": "Ah: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/14580/histctr\r\n\r\nAs explained by Gert. Which probably means that the site has reacted to user's search patterns and presented more relevant results based on the behavior in the training set.",
    "83682": "[quote=Gilberto Titericz Junior;83665]\r\n\r\nYour model and code seems to be ok!\r\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...\r\n\r\n\r\n\r\n[/quote]\r\n\r\nGilberto, is this CV with k-fold (k >1) or based in a time splitting with only a held data set?\r\n\r\nBy the way, can someone say me where I can fork Owen's script :-)",
    "83703": "[quote=José A. Guerrero;83682]\r\n\r\n[quote=Gilberto Titericz Junior;83665]\r\n\r\nYour model and code seems to be ok!\r\nMy CV lies around ~0.0306. So testset have different characteristics than trainset; Maybe time related or something else...\r\n\r\n\r\n\r\n[/quote]\r\n\r\nGilberto, is this CV with k-fold (k >1) or based in a time splitting with only a held data set?\r\n\r\nBy the way, can someone say me where I can fork Owen's script :-)\r\n\r\n\r\n\r\n\r\n\r\n[/quote]\r\n\r\n\r\nJosé: start here, it may help...\r\nhttps://github.com/owenzhang/kaggle-avazu\r\n\r\nMy cv is consistent with LB. Usually ~0.001 higher. And cv improvments reflects lb improvments.\r\n\r\nFor 0.04077 on lb, my cv is 0.04250",
    "83908": "The possible cause of the difference is a way the dataset was generated. As it is describes in \"Data Splitting\" section, the data is a result of the following steps:\r\n\r\n 1. select test time interval (May 12 - May 20)\r\n 2. take random sample of users having activity between May 12 and May 20\r\n 3. for each user select a test point (target impression), make the test set from these points\r\n 4. record all activities of the selected users in previous month and make the train set from it.\r\n\r\nThe problem is that now train set contains only those users that will be active at test time. The train set misses the users who get on the site, make one-three searches and leave it forever. And it seems that the latter ones behave differently than the 'regular' users with long search histories.\r\n\r\nIf you want to obtain a train set similar to test one, make it from the last user impressions before the test points. The HistCRT score on such set is 0.05747 which is quite closed to the 0.05717 of test set.",
    "84096": "I thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. \r\n\r\nI'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.\r\n\r\nWhat you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.\r\n\r\nSo my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.\r\n\r\nOther people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.",
    "84103": "By reading the Data Split section (last one) on the data page (https://www.kaggle.com/c/avito-context-ad-clicks/data). Its all there. Just make validation and training have the exact same relation that lb and the full train have. If I say more i'll just be handing the solution (actually I think I just did by saying so much)\r\n\r\n[quote=Remap on github;84096]\r\n\r\nI thought about this and did a lot of reprocessing to try to improve this. It is my primary concern at the moment, because in ML you have to have a mechanism for checking your results. Without such a representative set, you're only mucking about in a swamp like a moron. My submissions are all over the place, which tells me I haven't figured this out yet. Some submissions score 0.013 less, others score 0.005 less. \r\n\r\nI'm moving up on the leaderboard yes, although very slowly. What holds me back is that I don't have this representative set to test against, which allows me to investigate the changes I make to configurations quickly.\r\n\r\nWhat you're telling me is that the LB test set is not exactly representative of the training data we use, mostly because a host of users come in that click a lot more than historical users, so I suffer heavy penalties because of that. Unfortunately, I don't see how I can work around this besides figuring out how to assume more trigger happy clickers for unknown users for example, which is the wrong way around: I'd be adjusting my algorithm to the LB scores to calibrate my local results. If I change my algorithm I'm right back in the swamp.\r\n\r\nSo my attempt today was to use the last 1 or 2 days of training data as a validation set. I scored about 0.041 against that in 2 attempts with 2 different algorithms (my previous scores on cv were anything between 0.035-0.038, which eventually yielded ~0.013 delta against LB). The current delta against that is 0.005, which is still too high in my opinion, because it's the difference between knowing you're 0.046 or 0.041.\r\n\r\nOther people have reported on the forum they got pretty similar scores to the LB. My question here is a bit more theoretical... how do you figure out a proper CV set to use to test your data set against without testing this against the LB?.\r\n\r\n[/quote]",
    "84132": "Ok, so I process everything with flat files, because they're a lot faster to remodel data than SQL queries. It is indeed extremely important to pay very close attention to every detail here. I was trying to cut this short by simply tailing the end of the training file, but that doesn't create the correct distribution.\r\n\r\nPreviously the validation scored 0.041 on the very first run of 1M records. I think now I have the correct data to validate against. It doesn't change my position, but at least now I can play around more and get a feel for what changes in parameters really do.\r\n\r\nThank you very much!\r\n\r\n    Validation stage\r\n    ( (1000000, 0.043459011918673934, 0.043459011918673934))\r\n    ( (2000000, 0.04460827962764392, 0.0457575473366139))\r\n    ( (3000000, 0.04665980910677317, 0.05076286806503169))\r\n    ( (4000000, 0.04773306449442931, 0.05095283065739771))\r\n    ( (5000000, 0.04889447220230085, 0.053540103033787007))\r\n    ( (5459321, 0.04964257027849093))",
    "84157": "Looks great!\r\n\r\nA test run on 5% of the data:  0.047371\r\n\r\nLB submission: 0.04734"
  },
  "source": "meta"
}