{
  "id": 14516,
  "title": "Beating the Benchmark ;)",
  "url": "/competitions/avito-context-ad-clicks/discussion/14516",
  "author_name": "",
  "post_date": "2015-06-03T16:42:10.297Z",
  "votes": 12,
  "comment_count": 37,
  "views": 7266,
  "content": "<p>Based on tinrtgu's code:</p>\n<p>https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark</p>\n<p>All credits to&nbsp;<strong>tinrtgu </strong>!</p>",
  "messages": [
    {
      "id": "80793",
      "postDate": "06/03/2015 16:42:10",
      "content": "<p>Based on tinrtgu's code:</p>\n<p>https://www.kaggle.com/abhishek/avito-context-ad-clicks/beating-the-benchmark</p>\n<p>All credits to&nbsp;<strong>tinrtgu </strong>!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80799",
      "postDate": "06/03/2015 16:57:16",
      "content": "<p>Thanks for sharing!</p>\n<p>Haven't completed downloading yet. Seems a lot of work is needed for this one and ICDM&nbsp;one. Got too busy recently...</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80813",
      "postDate": "06/03/2015 19:42:16",
      "content": "<p>Nice one! Perhaps with subsampling you can make it finish?</p>\n<p>Tinrtgu, if you are out there, this benchmark does not excuse you from posting your own! :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80816",
      "postDate": "06/03/2015 19:57:15",
      "content": "<p>Will you have NULL targets where objecttype &lt;&gt; 3?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80897",
      "postDate": "06/04/2015 17:28:07",
      "content": "<p>is that tinrtgu's code from Avazu&nbsp;competition? :P</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80994",
      "postDate": "06/05/2015 10:43:00",
      "content": "<p>[quote=Pavitrakumar;80897]</p>\n<p>is that tinrtgu's code from Avazu&nbsp;competition? :P</p>\n<p>[/quote]</p>\n<p>yes</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "83220",
      "postDate": "07/02/2015 13:12:50",
      "content": "<p>Hi,</p>\n<p>can anyone give me a hint about the mathematics/algorithms behind that approach? I have no knowledge about Python, but have enough experience to read most programming languages. Nevertheless I cannot really grasp what is happening here ... except that apparently the algorithm depends only on the data in &quot;trainSearchStream.tsv&quot; ...</p>\n<p>Thanks,&nbsp;Olli</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "83234",
      "postDate": "07/02/2015 14:32:08",
      "content": "<p>[quote=Oliver Meyfarth;83220]</p>\n<p>Hi,</p>\n<p>can anyone give me a hint about the mathematics/algorithms behind that approach? I have no knowledge about Python, but have enough experience to read most programming languages. Nevertheless I cannot really grasp what is happening here ... except that apparently the algorithm depends only on the data in &quot;trainSearchStream.tsv&quot; ...</p>\n<p>Thanks,&nbsp;Olli</p>\n<p>[/quote]<br>It's &quot;Follow the Regularized Leader&quot; (FTRL) algorithm.<br>Originally developed by Google.<br>You can read more about it in the original paper - http://www.eecs.tufts.edu/~dsculley/papers/ad-click-prediction.pdf</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "83690",
      "postDate": "07/07/2015 19:04:13",
      "content": "<p>Hi,</p>\n\n<p>has anyone used this approach to improve upon the ~0.051 score it yields?</p>\n\n<p>I found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).</p>\n\n<p>So I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...</p>\n\n<p>Cheers</p>\n\n<p>Markus</p>",
      "rawMarkdown": "Hi,\r\n\r\nhas anyone used this approach to improve upon the ~0.051 score it yields?\r\n\r\nI found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).\r\n\r\nSo I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...\r\n\r\nCheers\r\n\r\nMarkus",
      "votes": null
    },
    {
      "id": "83702",
      "postDate": "07/07/2015 21:11:56",
      "content": "<p>[quote=MarDo;83690]</p>\n\n<p>Hi,</p>\n\n<p>has anyone used this approach to improve upon the ~0.051 score it yields?</p>\n\n<p>I found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).</p>\n\n<p>So I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...</p>\n\n<p>Cheers</p>\n\n<p>Markus</p>\n\n<p>[/quote]</p>\n\n<p>I really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.</p>",
      "rawMarkdown": "[quote=MarDo;83690]\r\n\r\nHi,\r\n\r\nhas anyone used this approach to improve upon the ~0.051 score it yields?\r\n\r\nI found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).\r\n\r\nSo I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...\r\n\r\nCheers\r\n\r\nMarkus\r\n\r\n\r\n\r\n[/quote]\r\n\r\nI really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.",
      "votes": null
    },
    {
      "id": "83736",
      "postDate": "07/08/2015 01:20:20",
      "content": "<p>[quote=MarDo;83690]</p>\n\n<p>Hi,</p>\n\n<p>has anyone used this approach to improve upon the ~0.051 score it yields?</p>\n\n<p>I found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).</p>\n\n<p>So I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...</p>\n\n<p>Cheers</p>\n\n<p>Markus</p>\n\n<p>[/quote]</p>\n\n<p>This is the winner model of several previous match</p>",
      "rawMarkdown": "[quote=MarDo;83690]\r\n\r\nHi,\r\n\r\nhas anyone used this approach to improve upon the ~0.051 score it yields?\r\n\r\nI found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).\r\n\r\nSo I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...\r\n\r\nCheers\r\n\r\nMarkus\r\n\r\n\r\n\r\n[/quote]\r\n\r\nThis is the winner model of several previous match",
      "votes": null
    },
    {
      "id": "83738",
      "postDate": "07/08/2015 01:52:49",
      "content": "<p>You guys are giving way too much away. But in this spirit of the legend of tinrtgu, I'd bet the next winner (this competition) will use extended tinrtgu as a large part of the ensemble to win yet again. His code is on par with VW. </p>",
      "rawMarkdown": "You guys are giving way too much away. But in this spirit of the legend of tinrtgu, I'd bet the next winner (this competition) will use extended tinrtgu as a large part of the ensemble to win yet again. His code is on par with VW.",
      "votes": null
    },
    {
      "id": "83767",
      "postDate": "07/08/2015 06:52:47",
      "content": "<p>Thanks for the replies so far. I wouldn't be afraid &quot;too much is given away&quot; here. While in other kaggle challenges I did ok, I am new to online learning an using this challenge mainly to learn. </p>\n\n<p>What I did as first steps was to adapt the code such that it considers more data from the training set as well as more features (from AdsInfo, e.g. categories &amp;  manually discretized price levels). Both did not improve the score at all. </p>\n\n<p>So either I am doing something wrong in this case (and I need to find out what that is) or it was just bad luck.</p>",
      "rawMarkdown": "Thanks for the replies so far. I wouldn't be afraid \"too much is given away\" here. While in other kaggle challenges I did ok, I am new to online learning an using this challenge mainly to learn. \r\n\r\nWhat I did as first steps was to adapt the code such that it considers more data from the training set as well as more features (from AdsInfo, e.g. categories &  manually discretized price levels). Both did not improve the score at all. \r\n\r\nSo either I am doing something wrong in this case (and I need to find out what that is) or it was just bad luck.",
      "votes": null
    },
    {
      "id": "83791",
      "postDate": "07/08/2015 13:47:50",
      "content": "<p>[quote=MarDo;83767]\nas well as more features (from AdsInfo, e.g. categories &amp;  manually discretized price levels)\n[/quote]\nThe ad price itself could be less important feature comparing with, say, relative price (when you analyzing price from the particular ad and nearest non-ads)</p>",
      "rawMarkdown": "[quote=MarDo;83767]\r\nas well as more features (from AdsInfo, e.g. categories &  manually discretized price levels)\r\n[/quote]\r\nThe ad price itself could be less important feature comparing with, say, relative price (when you analyzing price from the particular ad and nearest non-ads)",
      "votes": null
    },
    {
      "id": "83813",
      "postDate": "07/08/2015 17:12:30",
      "content": "<p>BTW, does anyone know how many LB this benchmark achieves?</p>",
      "rawMarkdown": "BTW, does anyone know how many LB this benchmark achieves?",
      "votes": null
    },
    {
      "id": "83864",
      "postDate": "07/08/2015 22:13:01",
      "content": "<p>[quote=Jiming Ye;83813]</p>\n\n<p>BTW, does anyone know how many LB this benchmark achieves?</p>\n\n<p>[/quote]\n0.051</p>",
      "rawMarkdown": "[quote=Jiming Ye;83813]\r\n\r\nBTW, does anyone know how many LB this benchmark achieves?\r\n\r\n[/quote]\r\n0.051",
      "votes": null
    },
    {
      "id": "83881",
      "postDate": "07/09/2015 01:50:26",
      "content": "<p>[quote=rcarson;83864]</p>\n\n<p>[quote=Jiming Ye;83813]</p>\n\n<p>BTW, does anyone know how many LB this benchmark achieves?</p>\n\n<p>[/quote]\n0.051</p>\n\n<p>[/quote]</p>\n\n<p>This is better than my current result, lol.</p>",
      "rawMarkdown": "[quote=rcarson;83864]\r\n\r\n[quote=Jiming Ye;83813]\r\n\r\nBTW, does anyone know how many LB this benchmark achieves?\r\n\r\n[/quote]\r\n0.051\r\n\r\n[/quote]\r\n\r\nThis is better than my current result, lol.",
      "votes": null
    },
    {
      "id": "84001",
      "postDate": "07/10/2015 04:39:10",
      "content": "<p>[quote=Leustagos;83702]</p>\n\n<p>I really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.</p>\n\n<p>[/quote]</p>\n\n<p>Leustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!</p>",
      "rawMarkdown": "[quote=Leustagos;83702]\r\n\r\nI really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.\r\n\r\n[/quote]\r\n\r\nLeustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!",
      "votes": null
    },
    {
      "id": "84025",
      "postDate": "07/10/2015 12:08:42",
      "content": "<p>FTRL is the algorithm, and i used the algorithm part of it. About the features, I surely did a whole bunch of them!\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.</p>\n\n<p>[quote=Vivek;84001]</p>\n\n<p>[quote=Leustagos;83702]</p>\n\n<p>I really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.</p>\n\n<p>[/quote]</p>\n\n<p>Leustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "FTRL is the algorithm, and i used the algorithm part of it. About the features, I surely did a whole bunch of them!\r\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.\r\n\r\n[quote=Vivek;84001]\r\n\r\n[quote=Leustagos;83702]\r\n\r\nI really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.\r\n\r\n[/quote]\r\n\r\nLeustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "84093",
      "postDate": "07/11/2015 01:49:25",
      "content": "<p>With alpha, beta, L1 and L2 of .... ? :)</p>",
      "rawMarkdown": "With alpha, beta, L1 and L2 of .... ? :)",
      "votes": null
    },
    {
      "id": "84098",
      "postDate": "07/11/2015 02:40:51",
      "content": "<p>@Leustagos, thank you!</p>\n\n<p>@Remap, I'm assuming they differ based on the choice of features, but the values used when only histCTR is a feature are given in the python script.</p>",
      "rawMarkdown": "Leustagos, thank you!\r\n\r\n@Remap, I'm assuming they differ based on the choice of features, but the values used when only histCTR is a feature are given in the python script.",
      "votes": null
    },
    {
      "id": "84100",
      "postDate": "07/11/2015 02:57:57",
      "content": "<p>The HistCTR is getting deleted prior to fitting the model: del line['HistCTR']</p>",
      "rawMarkdown": "The HistCTR is getting deleted prior to fitting the model: del line['HistCTR']",
      "votes": null
    },
    {
      "id": "84122",
      "postDate": "07/11/2015 09:16:21",
      "content": "<p>[quote=Leustagos;84025]</p>\n\n<p>I di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.</p>\n\n<p>[/quote]</p>\n\n<p>I would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. </p>\n\n<p>As I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). </p>\n\n<p>I tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. </p>\n\n<p>So I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....</p>",
      "rawMarkdown": "[quote=Leustagos;84025]\r\n\r\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.\r\n\r\n[/quote]\r\n\r\nI would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. \r\n\r\nAs I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). \r\n\r\nI tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. \r\n\r\nSo I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....",
      "votes": null
    },
    {
      "id": "84268",
      "postDate": "07/13/2015 11:11:24",
      "content": "<p>[quote=MarDo;84122]</p>\n\n<p>[quote=Leustagos;84025]</p>\n\n<p>I di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.</p>\n\n<p>[/quote]</p>\n\n<p>I would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. </p>\n\n<p>As I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). </p>\n\n<p>I tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. </p>\n\n<p>So I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....</p>\n\n<p>[/quote]</p>\n\n<p>I won`t tell the parameters, because it would be to much of a giveaway, right? But its not so hard to get them, and the method is pretty robust and gives similar results for many sets of parameters.\nI takes about 20 minutes for a full run, so one can tune the params pretty fast. When tunning the params I usually change just one at a time a see the direction of its improvment, them i do a kind of binary search with it. Cut it in half, or multiply by 2, them check in which interval it gives the most improvment and such. After tunning one I pursue the others. After doing it for all params, i go back at the first one and try again. The last step is doing a grid search using the range for each parameter that were promising.\nSo good luck tuning! you will learn far more if you do it yourself.</p>",
      "rawMarkdown": "[quote=MarDo;84122]\r\n\r\n[quote=Leustagos;84025]\r\n\r\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.\r\n\r\n[/quote]\r\n\r\nI would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. \r\n\r\nAs I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). \r\n\r\nI tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. \r\n\r\nSo I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....\r\n\r\n\r\n[/quote]\r\n\r\nI won`t tell the parameters, because it would be to much of a giveaway, right? But its not so hard to get them, and the method is pretty robust and gives similar results for many sets of parameters.\r\nI takes about 20 minutes for a full run, so one can tune the params pretty fast. When tunning the params I usually change just one at a time a see the direction of its improvment, them i do a kind of binary search with it. Cut it in half, or multiply by 2, them check in which interval it gives the most improvment and such. After tunning one I pursue the others. After doing it for all params, i go back at the first one and try again. The last step is doing a grid search using the range for each parameter that were promising.\r\nSo good luck tuning! you will learn far more if you do it yourself.",
      "votes": null
    },
    {
      "id": "85524",
      "postDate": "07/15/2015 12:33:47",
      "content": "<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>",
      "rawMarkdown": "It is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.",
      "votes": null
    },
    {
      "id": "85525",
      "postDate": "07/15/2015 12:49:23",
      "content": "<p>[quote=Bluefool;80816]</p>\n<p>Will you have NULL targets where objecttype &lt;&gt; 3?</p>\n<p>[/quote]</p>\n<p>For those of us coming to the competition late, is anyone able to help with a more basic question about ObjectType == 3 and the relationship to the target variable?</p>\n<p>Having read the data description, I thought the test data was only going to include contextual ads and therefore ObjectType 3, making the use of the other types for training an optional thing, seeing as there is no click info for them. But it looks like I must have misunderstood something as the test data has the full range of ObjectTypes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "85527",
      "postDate": "07/15/2015 13:06:20",
      "content": "<p>@leustagos: thank you so much. As I said earlier, it won't do bettttter than the class prior, but I could not figure out why and it was giving me a real headache! So this is quite a relief, although I maybe should have found it myself. Thanks again!</p>",
      "rawMarkdown": "leustagos: thank you so much. As I said earlier, it won't do bettttter than the class prior, but I could not figure out why and it was giving me a real headache! So this is quite a relief, although I maybe should have found it myself. Thanks again!",
      "votes": null
    },
    {
      "id": "85528",
      "postDate": "07/15/2015 13:08:24",
      "content": "<p>[quote=Leustagos;85524]</p>\n\n<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>\n\n<p>[/quote]</p>\n\n<p>Its not a bug and nothing wrong there. It was left for the users... ;)</p>",
      "rawMarkdown": "[quote=Leustagos;85524]\r\n\r\nIt is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.\r\n\r\n[/quote]\r\n\r\nIts not a bug and nothing wrong there. It was left for the users... ;)",
      "votes": null
    },
    {
      "id": "85530",
      "postDate": "07/15/2015 13:20:30",
      "content": "<p>[quote=Leustagos;85524]</p>\n\n<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>\n\n<p>[/quote]\nExactly Leustagos. I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. Not a lot of competitors on this competitions is probably a reason why this anomaly wasn't foreseen and reported by anyone else.</p>\n\n<p>Technically, it's not an error. Abhishek clearly named the variable &quot;bits&quot; and not the standard &quot;D&quot; as per VW/tinrtgu's code</p>",
      "rawMarkdown": "[quote=Leustagos;85524]\r\n\r\nIt is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.\r\n\r\n[/quote]\r\nExactly Leustagos. I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. Not a lot of competitors on this competitions is probably a reason why this anomaly wasn't foreseen and reported by anyone else.\r\n\r\nTechnically, it's not an error. Abhishek clearly named the variable \"bits\" and not the standard \"D\" as per VW/tinrtgu's code",
      "votes": null
    },
    {
      "id": "85532",
      "postDate": "07/15/2015 13:28:54",
      "content": "<p>So my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.</p>",
      "rawMarkdown": "So my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\r\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.",
      "votes": null
    },
    {
      "id": "85536",
      "postDate": "07/15/2015 14:00:57",
      "content": "<p>[quote=Remap on github;85532]</p>\n\n<p>So my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.</p>\n\n<p>[/quote]</p>\n\n<p>You should check 'bits' and 'learning_rate' parameters. The defaults are 18 and 0.5 respectively (I guess). They are different from what it is here. Also, by default, VW doesn't used vanilla SGD. It uses adaptive, invariant version. So, a difference is kinda expected. However, by playing with it for sometime, you should do equally well!</p>",
      "rawMarkdown": "[quote=Remap on github;85532]\r\n\r\nSo my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\r\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.\r\n\r\n[/quote]\r\n\r\nYou should check 'bits' and 'learning_rate' parameters. The defaults are 18 and 0.5 respectively (I guess). They are different from what it is here. Also, by default, VW doesn't used vanilla SGD. It uses adaptive, invariant version. So, a difference is kinda expected. However, by playing with it for sometime, you should do equally well!",
      "votes": null
    },
    {
      "id": "85545",
      "postDate": "07/15/2015 15:18:07",
      "content": "<p>I caught that one early, probably wouldn't even have gotten to my current position if I didn't catch that. \nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\nThat one is especially tricky, because the basic model works ok for 5-6 features or so, but as you add more features, the overall model produces more contributions to a prediction. I discovered severe issues in my latest training set productions, so can't say for sure ( my validation set is crapped out again ). But it's easy to see by adjusting it downwards, it makes the convergence much more aggressive, whilst setting it to something like sqrt(num_features) makes it feel right again.</p>",
      "rawMarkdown": "I caught that one early, probably wouldn't even have gotten to my current position if I didn't catch that. \r\nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\r\nThat one is especially tricky, because the basic model works ok for 5-6 features or so, but as you add more features, the overall model produces more contributions to a prediction. I discovered severe issues in my latest training set productions, so can't say for sure ( my validation set is crapped out again ). But it's easy to see by adjusting it downwards, it makes the convergence much more aggressive, whilst setting it to something like sqrt(num_features) makes it feel right again.",
      "votes": null
    },
    {
      "id": "85550",
      "postDate": "07/15/2015 15:44:53",
      "content": "<p>Definitely the name of the variable is wrong. If you let it as D, them it wouldn't confuse as much... By naming bit it should imply the number of bits used to build the hashing table... Just my two cents... :)\nAnyway, i used the original code because it had two way interactions included.</p>\n\n<p>[quote=Abhishek;85528]</p>\n\n<p>[quote=Leustagos;85524]</p>\n\n<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>\n\n<p>[/quote]</p>\n\n<p>Its not a bug and nothing wrong there. It was left for the users... ;)</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Definitely the name of the variable is wrong. If you let it as D, them it wouldn't confuse as much... By naming bit it should imply the number of bits used to build the hashing table... Just my two cents... :)\r\nAnyway, i used the original code because it had two way interactions included.\r\n\r\n[quote=Abhishek;85528]\r\n\r\n[quote=Leustagos;85524]\r\n\r\nIt is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.\r\n\r\n[/quote]\r\n\r\nIts not a bug and nothing wrong there. It was left for the users... ;)\r\n\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "85552",
      "postDate": "07/15/2015 15:56:24",
      "content": "<p>[quote=binga;85530]</p>\n\n<p>I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. </p>\n\n<p>[/quote]</p>\n\n<p>At least it wasn't too obvious also to others ;)</p>",
      "rawMarkdown": "[quote=binga;85530]\r\n\r\n I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. \r\n\r\n[/quote]\r\n\r\nAt least it wasn't too obvious also to others ;)",
      "votes": null
    },
    {
      "id": "85562",
      "postDate": "07/15/2015 17:25:42",
      "content": "<p>[quote=Remap on github;85545]\nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\n[/quote]</p>\n\n<p>So scratch that, it was silly. The paper even states how sigma is defined as 1/alpha, which means it's all correct.</p>\n\n<p>My dataset had some issues that caused some features to be random, so the algorithm had to hammer on the dataset a lot to get the real data to converge. My learning rate has shifted orders of magnitude now and I get decent convergence.</p>",
      "rawMarkdown": "[quote=Remap on github;85545]\r\nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\r\n[/quote]\r\n\r\nSo scratch that, it was silly. The paper even states how sigma is defined as 1/alpha, which means it's all correct.\r\n\r\nMy dataset had some issues that caused some features to be random, so the algorithm had to hammer on the dataset a lot to get the real data to converge. My learning rate has shifted orders of magnitude now and I get decent convergence.",
      "votes": null
    },
    {
      "id": "85585",
      "postDate": "07/15/2015 18:08:24",
      "content": "<p>&nbsp;The trainSearchStream contains all types of ads. however, the submission file only contains the ids of the contextual ads.&nbsp;</p>\n<p>[quote=lewis ml;85525]</p>\n<p>[quote=Bluefool;80816]</p>\n<p>Will you have NULL targets where objecttype &lt;&gt; 3?</p>\n<p>[/quote]</p>\n<p>For those of us coming to the competition late, is anyone able to help with a more basic question about ObjectType == 3 and the relationship to the target variable?</p>\n<p>Having read the data description, I thought the test data was only going to include contextual ads and therefore ObjectType 3, making the use of the other types for training an optional thing, seeing as there is no click info for them. But it looks like I must have misunderstood something as the test data has the full range of ObjectTypes.</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "85588",
      "postDate": "07/15/2015 18:40:12",
      "content": "<p>We were asked to predict the probability just for the Contextual ads (ObjectType = 3).</p>\n\n<p>Infact, if you have carefully observed, the train set doesn't have IsClick populated for other objecttypes. And only those IDs are present in sampleSubmission which are ObjectType = 3 in test set! So, that's some hint. You can easily reverse engineer.</p>\n\n<p>Approach - Subset the train and test by ObjectType=3, establish a robust Cross-validation process and you're good to go! You could start with the benchmark posted after subsetting the data.</p>\n\n<p>I hope this helps. Good luck!</p>",
      "rawMarkdown": "We were asked to predict the probability just for the Contextual ads (ObjectType = 3).\r\n\r\nInfact, if you have carefully observed, the train set doesn't have IsClick populated for other objecttypes. And only those IDs are present in sampleSubmission which are ObjectType = 3 in test set! So, that's some hint. You can easily reverse engineer.\r\n\r\nApproach - Subset the train and test by ObjectType=3, establish a robust Cross-validation process and you're good to go! You could start with the benchmark posted after subsetting the data.\r\n\r\nI hope this helps. Good luck!",
      "votes": null
    },
    {
      "id": "85593",
      "postDate": "07/15/2015 19:49:09",
      "content": "<p>Thanks for the replies and sorry for the premature question - I realised the sub is half the size of the test file as soon as I looked a little more closely at the data. And yes, now to investigate why a naive vowpal run gives a very low validation score (not a question, this time).</p>",
      "rawMarkdown": "Thanks for the replies and sorry for the premature question - I realised the sub is half the size of the test file as soon as I looked a little more closely at the data. And yes, now to investigate why a naive vowpal run gives a very low validation score (not a question, this time).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 80799,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "06/03/2015 16:57:16",
      "content": "<p>Thanks for sharing!</p>\n<p>Haven't completed downloading yet. Seems a lot of work is needed for this one and ICDM&nbsp;one. Got too busy recently...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80813,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "06/03/2015 19:42:16",
      "content": "<p>Nice one! Perhaps with subsampling you can make it finish?</p>\n<p>Tinrtgu, if you are out there, this benchmark does not excuse you from posting your own! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80816,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "06/03/2015 19:57:15",
      "content": "<p>Will you have NULL targets where objecttype &lt;&gt; 3?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80897,
      "author_name": "pavitrakumar78",
      "author_url": "",
      "post_date": "06/04/2015 17:28:07",
      "content": "<p>is that tinrtgu's code from Avazu&nbsp;competition? :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80994,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "06/05/2015 10:43:00",
      "content": "<p>[quote=Pavitrakumar;80897]</p>\n<p>is that tinrtgu's code from Avazu&nbsp;competition? :P</p>\n<p>[/quote]</p>\n<p>yes</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83220,
      "author_name": "olivermeyfarth",
      "author_url": "",
      "post_date": "07/02/2015 13:12:50",
      "content": "<p>Hi,</p>\n<p>can anyone give me a hint about the mathematics/algorithms behind that approach? I have no knowledge about Python, but have enough experience to read most programming languages. Nevertheless I cannot really grasp what is happening here ... except that apparently the algorithm depends only on the data in &quot;trainSearchStream.tsv&quot; ...</p>\n<p>Thanks,&nbsp;Olli</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83234,
      "author_name": "rootua",
      "author_url": "",
      "post_date": "07/02/2015 14:32:08",
      "content": "<p>[quote=Oliver Meyfarth;83220]</p>\n<p>Hi,</p>\n<p>can anyone give me a hint about the mathematics/algorithms behind that approach? I have no knowledge about Python, but have enough experience to read most programming languages. Nevertheless I cannot really grasp what is happening here ... except that apparently the algorithm depends only on the data in &quot;trainSearchStream.tsv&quot; ...</p>\n<p>Thanks,&nbsp;Olli</p>\n<p>[/quote]<br>It's &quot;Follow the Regularized Leader&quot; (FTRL) algorithm.<br>Originally developed by Google.<br>You can read more about it in the original paper - http://www.eecs.tufts.edu/~dsculley/papers/ad-click-prediction.pdf</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83690,
      "author_name": "mardoe",
      "author_url": "",
      "post_date": "07/07/2015 19:04:13",
      "content": "<p>Hi,</p>\n\n<p>has anyone used this approach to improve upon the ~0.051 score it yields?</p>\n\n<p>I found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).</p>\n\n<p>So I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...</p>\n\n<p>Cheers</p>\n\n<p>Markus</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83702,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/07/2015 21:11:56",
      "content": "<p>[quote=MarDo;83690]</p>\n\n<p>Hi,</p>\n\n<p>has anyone used this approach to improve upon the ~0.051 score it yields?</p>\n\n<p>I found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).</p>\n\n<p>So I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...</p>\n\n<p>Cheers</p>\n\n<p>Markus</p>\n\n<p>[/quote]</p>\n\n<p>I really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83736,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "07/08/2015 01:20:20",
      "content": "<p>[quote=MarDo;83690]</p>\n\n<p>Hi,</p>\n\n<p>has anyone used this approach to improve upon the ~0.051 score it yields?</p>\n\n<p>I found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).</p>\n\n<p>So I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...</p>\n\n<p>Cheers</p>\n\n<p>Markus</p>\n\n<p>[/quote]</p>\n\n<p>This is the winner model of several previous match</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83738,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "07/08/2015 01:52:49",
      "content": "<p>You guys are giving way too much away. But in this spirit of the legend of tinrtgu, I'd bet the next winner (this competition) will use extended tinrtgu as a large part of the ensemble to win yet again. His code is on par with VW. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83767,
      "author_name": "mardoe",
      "author_url": "",
      "post_date": "07/08/2015 06:52:47",
      "content": "<p>Thanks for the replies so far. I wouldn't be afraid &quot;too much is given away&quot; here. While in other kaggle challenges I did ok, I am new to online learning an using this challenge mainly to learn. </p>\n\n<p>What I did as first steps was to adapt the code such that it considers more data from the training set as well as more features (from AdsInfo, e.g. categories &amp;  manually discretized price levels). Both did not improve the score at all. </p>\n\n<p>So either I am doing something wrong in this case (and I need to find out what that is) or it was just bad luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83791,
      "author_name": "nightbat",
      "author_url": "",
      "post_date": "07/08/2015 13:47:50",
      "content": "<p>[quote=MarDo;83767]\nas well as more features (from AdsInfo, e.g. categories &amp;  manually discretized price levels)\n[/quote]\nThe ad price itself could be less important feature comparing with, say, relative price (when you analyzing price from the particular ad and nearest non-ads)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83813,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "07/08/2015 17:12:30",
      "content": "<p>BTW, does anyone know how many LB this benchmark achieves?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83864,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "07/08/2015 22:13:01",
      "content": "<p>[quote=Jiming Ye;83813]</p>\n\n<p>BTW, does anyone know how many LB this benchmark achieves?</p>\n\n<p>[/quote]\n0.051</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 83881,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "07/09/2015 01:50:26",
      "content": "<p>[quote=rcarson;83864]</p>\n\n<p>[quote=Jiming Ye;83813]</p>\n\n<p>BTW, does anyone know how many LB this benchmark achieves?</p>\n\n<p>[/quote]\n0.051</p>\n\n<p>[/quote]</p>\n\n<p>This is better than my current result, lol.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84001,
      "author_name": "vman049",
      "author_url": "",
      "post_date": "07/10/2015 04:39:10",
      "content": "<p>[quote=Leustagos;83702]</p>\n\n<p>I really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.</p>\n\n<p>[/quote]</p>\n\n<p>Leustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84025,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/10/2015 12:08:42",
      "content": "<p>FTRL is the algorithm, and i used the algorithm part of it. About the features, I surely did a whole bunch of them!\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.</p>\n\n<p>[quote=Vivek;84001]</p>\n\n<p>[quote=Leustagos;83702]</p>\n\n<p>I really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.</p>\n\n<p>[/quote]</p>\n\n<p>Leustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84093,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/11/2015 01:49:25",
      "content": "<p>With alpha, beta, L1 and L2 of .... ? :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84098,
      "author_name": "vman049",
      "author_url": "",
      "post_date": "07/11/2015 02:40:51",
      "content": "<p>@Leustagos, thank you!</p>\n\n<p>@Remap, I'm assuming they differ based on the choice of features, but the values used when only histCTR is a feature are given in the python script.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84100,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/11/2015 02:57:57",
      "content": "<p>The HistCTR is getting deleted prior to fitting the model: del line['HistCTR']</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84122,
      "author_name": "mardoe",
      "author_url": "",
      "post_date": "07/11/2015 09:16:21",
      "content": "<p>[quote=Leustagos;84025]</p>\n\n<p>I di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.</p>\n\n<p>[/quote]</p>\n\n<p>I would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. </p>\n\n<p>As I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). </p>\n\n<p>I tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. </p>\n\n<p>So I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 84268,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/13/2015 11:11:24",
      "content": "<p>[quote=MarDo;84122]</p>\n\n<p>[quote=Leustagos;84025]</p>\n\n<p>I di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.</p>\n\n<p>[/quote]</p>\n\n<p>I would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. </p>\n\n<p>As I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). </p>\n\n<p>I tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. </p>\n\n<p>So I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....</p>\n\n<p>[/quote]</p>\n\n<p>I won`t tell the parameters, because it would be to much of a giveaway, right? But its not so hard to get them, and the method is pretty robust and gives similar results for many sets of parameters.\nI takes about 20 minutes for a full run, so one can tune the params pretty fast. When tunning the params I usually change just one at a time a see the direction of its improvment, them i do a kind of binary search with it. Cut it in half, or multiply by 2, them check in which interval it gives the most improvment and such. After tunning one I pursue the others. After doing it for all params, i go back at the first one and try again. The last step is doing a grid search using the range for each parameter that were promising.\nSo good luck tuning! you will learn far more if you do it yourself.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85524,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/15/2015 12:33:47",
      "content": "<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85525,
      "author_name": "cartographic",
      "author_url": "",
      "post_date": "07/15/2015 12:49:23",
      "content": "<p>[quote=Bluefool;80816]</p>\n<p>Will you have NULL targets where objecttype &lt;&gt; 3?</p>\n<p>[/quote]</p>\n<p>For those of us coming to the competition late, is anyone able to help with a more basic question about ObjectType == 3 and the relationship to the target variable?</p>\n<p>Having read the data description, I thought the test data was only going to include contextual ads and therefore ObjectType 3, making the use of the other types for training an optional thing, seeing as there is no click info for them. But it looks like I must have misunderstood something as the test data has the full range of ObjectTypes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85527,
      "author_name": "mardoe",
      "author_url": "",
      "post_date": "07/15/2015 13:06:20",
      "content": "<p>@leustagos: thank you so much. As I said earlier, it won't do bettttter than the class prior, but I could not figure out why and it was giving me a real headache! So this is quite a relief, although I maybe should have found it myself. Thanks again!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85528,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "07/15/2015 13:08:24",
      "content": "<p>[quote=Leustagos;85524]</p>\n\n<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>\n\n<p>[/quote]</p>\n\n<p>Its not a bug and nothing wrong there. It was left for the users... ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85530,
      "author_name": "phanisrikanth",
      "author_url": "",
      "post_date": "07/15/2015 13:20:30",
      "content": "<p>[quote=Leustagos;85524]</p>\n\n<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>\n\n<p>[/quote]\nExactly Leustagos. I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. Not a lot of competitors on this competitions is probably a reason why this anomaly wasn't foreseen and reported by anyone else.</p>\n\n<p>Technically, it's not an error. Abhishek clearly named the variable &quot;bits&quot; and not the standard &quot;D&quot; as per VW/tinrtgu's code</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85532,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/15/2015 13:28:54",
      "content": "<p>So my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85536,
      "author_name": "phanisrikanth",
      "author_url": "",
      "post_date": "07/15/2015 14:00:57",
      "content": "<p>[quote=Remap on github;85532]</p>\n\n<p>So my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.</p>\n\n<p>[/quote]</p>\n\n<p>You should check 'bits' and 'learning_rate' parameters. The defaults are 18 and 0.5 respectively (I guess). They are different from what it is here. Also, by default, VW doesn't used vanilla SGD. It uses adaptive, invariant version. So, a difference is kinda expected. However, by playing with it for sometime, you should do equally well!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85545,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/15/2015 15:18:07",
      "content": "<p>I caught that one early, probably wouldn't even have gotten to my current position if I didn't catch that. \nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\nThat one is especially tricky, because the basic model works ok for 5-6 features or so, but as you add more features, the overall model produces more contributions to a prediction. I discovered severe issues in my latest training set productions, so can't say for sure ( my validation set is crapped out again ). But it's easy to see by adjusting it downwards, it makes the convergence much more aggressive, whilst setting it to something like sqrt(num_features) makes it feel right again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85550,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/15/2015 15:44:53",
      "content": "<p>Definitely the name of the variable is wrong. If you let it as D, them it wouldn't confuse as much... By naming bit it should imply the number of bits used to build the hashing table... Just my two cents... :)\nAnyway, i used the original code because it had two way interactions included.</p>\n\n<p>[quote=Abhishek;85528]</p>\n\n<p>[quote=Leustagos;85524]</p>\n\n<p>It is worth mentioning that in this benchmark bits is defined wrongly.\ninstead of using <strong>bits = 20, we should use bit = 2**20, or 2**24</strong> should be a better starting point.\nThe its defined right now won't do much better than defining a single average.</p>\n\n<p>[/quote]</p>\n\n<p>Its not a bug and nothing wrong there. It was left for the users... ;)</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85552,
      "author_name": "mardoe",
      "author_url": "",
      "post_date": "07/15/2015 15:56:24",
      "content": "<p>[quote=binga;85530]</p>\n\n<p>I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. </p>\n\n<p>[/quote]</p>\n\n<p>At least it wasn't too obvious also to others ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85562,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/15/2015 17:25:42",
      "content": "<p>[quote=Remap on github;85545]\nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\n[/quote]</p>\n\n<p>So scratch that, it was silly. The paper even states how sigma is defined as 1/alpha, which means it's all correct.</p>\n\n<p>My dataset had some issues that caused some features to be random, so the algorithm had to hammer on the dataset a lot to get the real data to converge. My learning rate has shifted orders of magnitude now and I get decent convergence.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85585,
      "author_name": "mas313",
      "author_url": "",
      "post_date": "07/15/2015 18:08:24",
      "content": "<p>&nbsp;The trainSearchStream contains all types of ads. however, the submission file only contains the ids of the contextual ads.&nbsp;</p>\n<p>[quote=lewis ml;85525]</p>\n<p>[quote=Bluefool;80816]</p>\n<p>Will you have NULL targets where objecttype &lt;&gt; 3?</p>\n<p>[/quote]</p>\n<p>For those of us coming to the competition late, is anyone able to help with a more basic question about ObjectType == 3 and the relationship to the target variable?</p>\n<p>Having read the data description, I thought the test data was only going to include contextual ads and therefore ObjectType 3, making the use of the other types for training an optional thing, seeing as there is no click info for them. But it looks like I must have misunderstood something as the test data has the full range of ObjectTypes.</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85588,
      "author_name": "phanisrikanth",
      "author_url": "",
      "post_date": "07/15/2015 18:40:12",
      "content": "<p>We were asked to predict the probability just for the Contextual ads (ObjectType = 3).</p>\n\n<p>Infact, if you have carefully observed, the train set doesn't have IsClick populated for other objecttypes. And only those IDs are present in sampleSubmission which are ObjectType = 3 in test set! So, that's some hint. You can easily reverse engineer.</p>\n\n<p>Approach - Subset the train and test by ObjectType=3, establish a robust Cross-validation process and you're good to go! You could start with the benchmark posted after subsetting the data.</p>\n\n<p>I hope this helps. Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85593,
      "author_name": "cartographic",
      "author_url": "",
      "post_date": "07/15/2015 19:49:09",
      "content": "<p>Thanks for the replies and sorry for the premature question - I realised the sub is half the size of the test file as soon as I looked a little more closely at the data. And yes, now to investigate why a naive vowpal run gives a very low validation score (not a question, this time).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "80793": "",
    "80799": "",
    "80813": "",
    "80816": "",
    "80897": "",
    "80994": "",
    "83220": "",
    "83234": "",
    "83690": "Hi,\r\n\r\nhas anyone used this approach to improve upon the ~0.051 score it yields?\r\n\r\nI found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).\r\n\r\nSo I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...\r\n\r\nCheers\r\n\r\nMarkus",
    "83702": "[quote=MarDo;83690]\r\n\r\nHi,\r\n\r\nhas anyone used this approach to improve upon the ~0.051 score it yields?\r\n\r\nI found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).\r\n\r\nSo I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...\r\n\r\nCheers\r\n\r\nMarkus\r\n\r\n\r\n\r\n[/quote]\r\n\r\nI really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.",
    "83736": "[quote=MarDo;83690]\r\n\r\nHi,\r\n\r\nhas anyone used this approach to improve upon the ~0.051 score it yields?\r\n\r\nI found that a similar score can already be achieved by just submitting the prior (which should actually be a default benchmark for suchlike challenges :) ).\r\n\r\nSo I am not yet convinced about the advantage of this approach, at least not from what I have seen in this thread...\r\n\r\nCheers\r\n\r\nMarkus\r\n\r\n\r\n\r\n[/quote]\r\n\r\nThis is the winner model of several previous match",
    "83738": "You guys are giving way too much away. But in this spirit of the legend of tinrtgu, I'd bet the next winner (this competition) will use extended tinrtgu as a large part of the ensemble to win yet again. His code is on par with VW.",
    "83767": "Thanks for the replies so far. I wouldn't be afraid \"too much is given away\" here. While in other kaggle challenges I did ok, I am new to online learning an using this challenge mainly to learn. \r\n\r\nWhat I did as first steps was to adapt the code such that it considers more data from the training set as well as more features (from AdsInfo, e.g. categories &  manually discretized price levels). Both did not improve the score at all. \r\n\r\nSo either I am doing something wrong in this case (and I need to find out what that is) or it was just bad luck.",
    "83791": "[quote=MarDo;83767]\r\nas well as more features (from AdsInfo, e.g. categories &  manually discretized price levels)\r\n[/quote]\r\nThe ad price itself could be less important feature comparing with, say, relative price (when you analyzing price from the particular ad and nearest non-ads)",
    "83813": "BTW, does anyone know how many LB this benchmark achieves?",
    "83864": "[quote=Jiming Ye;83813]\r\n\r\nBTW, does anyone know how many LB this benchmark achieves?\r\n\r\n[/quote]\r\n0.051",
    "83881": "[quote=rcarson;83864]\r\n\r\n[quote=Jiming Ye;83813]\r\n\r\nBTW, does anyone know how many LB this benchmark achieves?\r\n\r\n[/quote]\r\n0.051\r\n\r\n[/quote]\r\n\r\nThis is better than my current result, lol.",
    "84001": "[quote=Leustagos;83702]\r\n\r\nI really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.\r\n\r\n[/quote]\r\n\r\nLeustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!",
    "84025": "FTRL is the algorithm, and i used the algorithm part of it. About the features, I surely did a whole bunch of them!\r\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.\r\n\r\n[quote=Vivek;84001]\r\n\r\n[quote=Leustagos;83702]\r\n\r\nI really recommend you to look again. Also you can check previous winners posts and see the usefulness of FTRL. By the way we did beat 0.051 with ftrl by a long margin.\r\n\r\n[/quote]\r\n\r\nLeustagos, please forgive me if I'm prying too much, but do you mean that you used (a modified) FTRL to beat 0.051 (perhaps with more features than just histCTR) or simply that you beat FTRL using another method of your own? Thanks!\r\n\r\n[/quote]",
    "84093": "With alpha, beta, L1 and L2 of .... ? :)",
    "84098": "Leustagos, thank you!\r\n\r\n@Remap, I'm assuming they differ based on the choice of features, but the values used when only histCTR is a feature are given in the python script.",
    "84100": "The HistCTR is getting deleted prior to fitting the model: del line['HistCTR']",
    "84122": "[quote=Leustagos;84025]\r\n\r\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.\r\n\r\n[/quote]\r\n\r\nI would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. \r\n\r\nAs I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). \r\n\r\nI tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. \r\n\r\nSo I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....",
    "84268": "[quote=MarDo;84122]\r\n\r\n[quote=Leustagos;84025]\r\n\r\nI di remenber that FTRL using only AdID and position as features gets 0.048X if trained on the whole dataset.\r\n\r\n[/quote]\r\n\r\nI would also be interested in the parameters, however IMO you don't have to reveal them here. Instead of gaining some positions in the leaderboard by reusing the findings of others, I would be much more interested in the general methodology how one determines which parameters to set. \r\n\r\nAs I said before, using the above provided code and running it on the whole dataset is not better than just submitting the prior probability of the target class (or I did something wrong). \r\n\r\nI tried different parameter constellations (e.g. from previous challenges using FTRL) - no change of result. \r\n\r\nSo I am aware of general grid search (and maybe more advanced approaches for parameter optimization), but here I am not quite sure where to start and where to stop....\r\n\r\n\r\n[/quote]\r\n\r\nI won`t tell the parameters, because it would be to much of a giveaway, right? But its not so hard to get them, and the method is pretty robust and gives similar results for many sets of parameters.\r\nI takes about 20 minutes for a full run, so one can tune the params pretty fast. When tunning the params I usually change just one at a time a see the direction of its improvment, them i do a kind of binary search with it. Cut it in half, or multiply by 2, them check in which interval it gives the most improvment and such. After tunning one I pursue the others. After doing it for all params, i go back at the first one and try again. The last step is doing a grid search using the range for each parameter that were promising.\r\nSo good luck tuning! you will learn far more if you do it yourself.",
    "85524": "It is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.",
    "85525": "",
    "85527": "leustagos: thank you so much. As I said earlier, it won't do bettttter than the class prior, but I could not figure out why and it was giving me a real headache! So this is quite a relief, although I maybe should have found it myself. Thanks again!",
    "85528": "[quote=Leustagos;85524]\r\n\r\nIt is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.\r\n\r\n[/quote]\r\n\r\nIts not a bug and nothing wrong there. It was left for the users... ;)",
    "85530": "[quote=Leustagos;85524]\r\n\r\nIt is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.\r\n\r\n[/quote]\r\nExactly Leustagos. I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. Not a lot of competitors on this competitions is probably a reason why this anomaly wasn't foreseen and reported by anyone else.\r\n\r\nTechnically, it's not an error. Abhishek clearly named the variable \"bits\" and not the standard \"D\" as per VW/tinrtgu's code",
    "85532": "So my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\r\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.",
    "85536": "[quote=Remap on github;85532]\r\n\r\nSo my submissions until now used FTRL. Is it possible to use vowpal wabbit to beat the benchmark the same way. I noticed it has --ftrl flags in there, but I'm not sure if the rest of the algorithm is the same. I did a test pass, but the results from vw are significantly worse. Should one not be getting the same results more or less, running against the exact same test data with the same parameters?\r\nI also ran vw without many parameters for tweaking and that didn't produce exactly stunning results on my validation set.\r\n\r\n[/quote]\r\n\r\nYou should check 'bits' and 'learning_rate' parameters. The defaults are 18 and 0.5 respectively (I guess). They are different from what it is here. Also, by default, VW doesn't used vanilla SGD. It uses adaptive, invariant version. So, a difference is kinda expected. However, by playing with it for sometime, you should do equally well!",
    "85545": "I caught that one early, probably wouldn't even have gotten to my current position if I didn't catch that. \r\nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\r\nThat one is especially tricky, because the basic model works ok for 5-6 features or so, but as you add more features, the overall model produces more contributions to a prediction. I discovered severe issues in my latest training set productions, so can't say for sure ( my validation set is crapped out again ). But it's easy to see by adjusting it downwards, it makes the convergence much more aggressive, whilst setting it to something like sqrt(num_features) makes it feel right again.",
    "85550": "Definitely the name of the variable is wrong. If you let it as D, them it wouldn't confuse as much... By naming bit it should imply the number of bits used to build the hashing table... Just my two cents... :)\r\nAnyway, i used the original code because it had two way interactions included.\r\n\r\n[quote=Abhishek;85528]\r\n\r\n[quote=Leustagos;85524]\r\n\r\nIt is worth mentioning that in this benchmark bits is defined wrongly.\r\ninstead of using **bits = 20, we should use bit = 2**20, or 2**24** should be a better starting point.\r\nThe its defined right now won't do much better than defining a single average.\r\n\r\n[/quote]\r\n\r\nIts not a bug and nothing wrong there. It was left for the users... ;)\r\n\r\n\r\n[/quote]",
    "85552": "[quote=binga;85530]\r\n\r\n I broke my head for a good couple of days before figuring out the prime difference between the legendary tinrtgu's code and the benchmark posted here! I could finally catch it. \r\n\r\n[/quote]\r\n\r\nAt least it wasn't too obvious also to others ;)",
    "85562": "[quote=Remap on github;85545]\r\nThe other thing is the alpha learning parameter.  I think it's the other way around in the sense that you must increase it to decrease the effect?\r\n[/quote]\r\n\r\nSo scratch that, it was silly. The paper even states how sigma is defined as 1/alpha, which means it's all correct.\r\n\r\nMy dataset had some issues that caused some features to be random, so the algorithm had to hammer on the dataset a lot to get the real data to converge. My learning rate has shifted orders of magnitude now and I get decent convergence.",
    "85585": "",
    "85588": "We were asked to predict the probability just for the Contextual ads (ObjectType = 3).\r\n\r\nInfact, if you have carefully observed, the train set doesn't have IsClick populated for other objecttypes. And only those IDs are present in sampleSubmission which are ObjectType = 3 in test set! So, that's some hint. You can easily reverse engineer.\r\n\r\nApproach - Subset the train and test by ObjectType=3, establish a robust Cross-validation process and you're good to go! You could start with the benchmark posted after subsetting the data.\r\n\r\nI hope this helps. Good luck!",
    "85593": "Thanks for the replies and sorry for the premature question - I realised the sub is half the size of the test file as soon as I looked a little more closely at the data. And yes, now to investigate why a naive vowpal run gives a very low validation score (not a question, this time)."
  },
  "source": "meta"
}