{
  "id": 15606,
  "title": "Congratulations!",
  "url": "/competitions/avito-context-ad-clicks/discussion/15606",
  "author_name": "",
  "post_date": "2015-07-29T00:28:42.187Z",
  "votes": 3,
  "comment_count": 19,
  "views": 5736,
  "content": "<p>Great contest! Congratulations to top players! To be honest I studied several previous winning solutions but still cannot figure out this one. I'm especially interested in the sampling and quick validation strategies. Thank you all in advance.</p>",
  "messages": [
    {
      "id": "87358",
      "postDate": "07/29/2015 00:28:42",
      "content": "<p>Great contest! Congratulations to top players! To be honest I studied several previous winning solutions but still cannot figure out this one. I'm especially interested in the sampling and quick validation strategies. Thank you all in advance.</p>",
      "rawMarkdown": "Great contest! Congratulations to top players! To be honest I studied several previous winning solutions but still cannot figure out this one. I'm especially interested in the sampling and quick validation strategies. Thank you all in advance.",
      "votes": null
    },
    {
      "id": "87362",
      "postDate": "07/29/2015 00:45:33",
      "content": "<p>Congratulations to the top 10 and all other competitors that made this exciting.  We tried pushing for top 10 but couldn't make it at the end.  In fact, our single Vowpal Wabbit model would have still placed us in 11th still.  2 important things we did in this competition:</p>\n\n<ol>\n<li>Sorted users' searches by time so that each place (e.g., last search by user, second to last search by user, etc) were stored in separate training files.  Then trained on the first 9 of the last 10 searches in training (in the order) and validating on the last search before test.  This reduced the training to only ~30 Million records.  This correlated near perfectly with the Public LB.  Final submission was trained on only the last 99 searches in the training set (~112M records).</li>\n<li>Number of ads, and number of types of ads (Types 1, 2, &amp; 3) on the search provided one of the biggest boosts.</li>\n</ol>\n\n<p>We should have used the VisitStreams and PhoneRequestStream, but ran out of time.</p>",
      "rawMarkdown": "Congratulations to the top 10 and all other competitors that made this exciting.  We tried pushing for top 10 but couldn't make it at the end.  In fact, our single Vowpal Wabbit model would have still placed us in 11th still.  2 important things we did in this competition:\r\n\r\n1. Sorted users' searches by time so that each place (e.g., last search by user, second to last search by user, etc) were stored in separate training files.  Then trained on the first 9 of the last 10 searches in training (in the order) and validating on the last search before test.  This reduced the training to only ~30 Million records.  This correlated near perfectly with the Public LB.  Final submission was trained on only the last 99 searches in the training set (~112M records).\r\n2. Number of ads, and number of types of ads (Types 1, 2, & 3) on the search provided one of the biggest boosts.\r\n\r\nWe should have used the VisitStreams and PhoneRequestStream, but ran out of time.",
      "votes": null
    },
    {
      "id": "87372",
      "postDate": "07/29/2015 01:49:59",
      "content": "<p>I couldn't figure this competition out either. I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? </p>\n\n<p>I did, however, use all the data including both the phone and visits stream. I also engineered features including bag of Russian words (attached). I'm not sure if it was a correct way of going about it, but it seemed to give a small boost.</p>",
      "rawMarkdown": "I couldn't figure this competition out either. I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? \r\n\r\nI did, however, use all the data including both the phone and visits stream. I also engineered features including bag of Russian words (attached). I'm not sure if it was a correct way of going about it, but it seemed to give a small boost.",
      "votes": null
    },
    {
      "id": "87373",
      "postDate": "07/29/2015 02:10:21",
      "content": "<p>[quote=Mike Kim;87372]</p>\n\n<p>I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? </p>\n\n<p>[/quote]\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:</p>\n\n<pre><code>vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0\n</code></pre>",
      "rawMarkdown": "[quote=Mike Kim;87372]\r\n\r\nI couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? \r\n\r\n[/quote]\r\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:\r\n\r\n    vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0",
      "votes": null
    },
    {
      "id": "87374",
      "postDate": "07/29/2015 02:21:19",
      "content": "<p>Congratulations to the winners,  I'm very lucky in the private leaderboard, I spent a lot of time pushing myself in top2, but when I thought a idea to improve my scores, I always saw that Owen and Dmitry &amp; Leustagos improve their scores again, this is a really hard competition.</p>\n\n<p>I list some key points.</p>\n\n<p>(1). negative sampling(0.1) to reduce data size(about 20M records).</p>\n\n<p>(2). a lot of features, like number of ads, rank by price for every ads, position type(all ads' position in one tuple, then use hash value as a category feature), show count and click count for every users,  and so on. But I also don't use VisitStreams and PhoneRequestStream, maybe I should try it.\nBy the way, I just test new features in first 10M records of trainSearchStream.tsv, it looks like very effective.</p>\n\n<p>(3). ensemble. FFM, FM, GBDT as base models, then use NN and GBDT to combine it. Finally, I average my best 3 submissions. </p>\n\n<p>best model is a FFM model, it give me 0.04061 in public leaderboard.</p>",
      "rawMarkdown": "Congratulations to the winners,  I'm very lucky in the private leaderboard, I spent a lot of time pushing myself in top2, but when I thought a idea to improve my scores, I always saw that Owen and Dmitry & Leustagos improve their scores again, this is a really hard competition.\r\n\r\nI list some key points.\r\n\r\n(1). negative sampling(0.1) to reduce data size(about 20M records).\r\n\r\n(2). a lot of features, like number of ads, rank by price for every ads, position type(all ads' position in one tuple, then use hash value as a category feature), show count and click count for every users,  and so on. But I also don't use VisitStreams and PhoneRequestStream, maybe I should try it.\r\nBy the way, I just test new features in first 10M records of trainSearchStream.tsv, it looks like very effective.\r\n\r\n(3). ensemble. FFM, FM, GBDT as base models, then use NN and GBDT to combine it. Finally, I average my best 3 submissions. \r\n\r\nbest model is a FFM model, it give me 0.04061 in public leaderboard.",
      "votes": null
    },
    {
      "id": "87377",
      "postDate": "07/29/2015 02:44:17",
      "content": "<p>I don't think that's the issue since I hashed tricked every single string in Python when parsing the file for VW format. Then again my parsing script could have been buggy. I kept getting saturated predictions of exactly 0 or 1 or (-50, or 50 when raw?) or a seg fault within VW. I think it's a system problem since I recently got a new computer (64GB ram which seems like not enough) and how to reinstall VW and everything else.\n[quote=David Shinn;87373]</p>\n\n<p>[quote=Mike Kim;87372]</p>\n\n<p>I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? </p>\n\n<p>[/quote]\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:</p>\n\n<pre><code>vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0\n</code></pre>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I don't think that's the issue since I hashed tricked every single string in Python when parsing the file for VW format. Then again my parsing script could have been buggy. I kept getting saturated predictions of exactly 0 or 1 or (-50, or 50 when raw?) or a seg fault within VW. I think it's a system problem since I recently got a new computer (64GB ram which seems like not enough) and how to reinstall VW and everything else.\r\n[quote=David Shinn;87373]\r\n\r\n[quote=Mike Kim;87372]\r\n\r\nI couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? \r\n\r\n[/quote]\r\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:\r\n\r\n    vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "87379",
      "postDate": "07/29/2015 03:14:46",
      "content": "<p>Congratulations to the winners!</p>\n\n<p>I spent at least half of the time trying to employ VisitsStream and RequestStream in my model (FFM).   I tried a number of straightforward features like AD words, stemmed words, category, time etc. But each time I was getting only marginal improvement. The last and most promising idea was to compose  'sentences' from the titles of ADs consequently viewed by a user and train word2vec on it. Then it can be used in many ways to generate features based on distance between ADs or query to AD or ads classification etc.  Unfortunately I ran out of time implementing this but preliminary tests on sample showed impressive improvement.</p>\n\n<p>If somebody successfully used personalization data I would be happy if you share a way you did it. </p>",
      "rawMarkdown": "Congratulations to the winners!\r\n\r\nI spent at least half of the time trying to employ VisitsStream and RequestStream in my model (FFM).   I tried a number of straightforward features like AD words, stemmed words, category, time etc. But each time I was getting only marginal improvement. The last and most promising idea was to compose  'sentences' from the titles of ADs consequently viewed by a user and train word2vec on it. Then it can be used in many ways to generate features based on distance between ADs or query to AD or ads classification etc.  Unfortunately I ran out of time implementing this but preliminary tests on sample showed impressive improvement.\r\n\r\nIf somebody successfully used personalization data I would be happy if you share a way you did it.",
      "votes": null
    },
    {
      "id": "87405",
      "postDate": "07/29/2015 11:33:31",
      "content": "<p>Congratulation to winners and all top teams! Well done!</p>\n\n<p>This was a very interesting competition! Improvements on validation set reflected the same improvements on LB. Our validation set was the same as that described here: <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15367/proper-validation-set/86077#post86077\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15367/proper-validation-set/86077#post86077</a>.\nAnd it works perfectly:)</p>\n\n<p>Our approach:</p>\n\n<p>1) My single ftrl model with basic features: AdID, UserID, IPID, Position, Price, &#8230; + the second- and the third-order interactions between some of this features + scores similarity between SearchQuery and Title and SearchParams and Params. It gets Public LB score below 0.044.</p>\n\n<p>2) Adding to this model only one new feature PositionFactor from my teammate Alexander. It gave the incredible increments and the model scored  ~ 0.0418 on Public LB. </p>\n\n<p>PositionFactor = hash( [Position 1_place:ObjectType 2_place:ObjectType 6_place:ObjectType 7_place:ObjectType 8_place:ObjectType]) , where \nPosition is the position of the given context ad in search result page.\nk_place:ObjectType is the position and type others ads in search result page (if there are data about them).</p>\n\n<p>3) Alexander's models vw+xgb and vw+rf give him ~ 0.0424 on Public LB</p>\n\n<p>4) After tuning hyperparameters and finding the weights for linear combination of our solutions we achieved 0.04137 on Public LB. </p>",
      "rawMarkdown": "Congratulation to winners and all top teams! Well done!\r\n\r\nThis was a very interesting competition! Improvements on validation set reflected the same improvements on LB. Our validation set was the same as that described here: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15367/proper-validation-set/86077#post86077.\r\nAnd it works perfectly:)\r\n\r\nOur approach:\r\n\r\n1) My single ftrl model with basic features: AdID, UserID, IPID, Position, Price, … + the second- and the third-order interactions between some of this features + scores similarity between SearchQuery and Title and SearchParams and Params. It gets Public LB score below 0.044.\r\n\r\n2) Adding to this model only one new feature PositionFactor from my teammate Alexander. It gave the incredible increments and the model scored  ~ 0.0418 on Public LB. \r\n\r\nPositionFactor = hash( [Position 1_place:ObjectType 2_place:ObjectType 6_place:ObjectType 7_place:ObjectType 8_place:ObjectType]) , where \r\nPosition is the position of the given context ad in search result page.\r\nk_place:ObjectType is the position and type others ads in search result page (if there are data about them).\r\n\r\n3) Alexander's models vw+xgb and vw+rf give him ~ 0.0424 on Public LB\r\n\r\n4) After tuning hyperparameters and finding the weights for linear combination of our solutions we achieved 0.04137 on Public LB.",
      "votes": null
    },
    {
      "id": "87421",
      "postDate": "07/29/2015 13:09:57",
      "content": "<p>Congratulations to the winners!!! That was a unique big data competition!</p>\n\n<p>Our aproach is based in two algos: FTRL and XGBoost. Basically we build 3 models  for each algo, then found some weights to combine then all.</p>\n\n<p>FTRL was built over all dataset:\nModel1 -sorted randomly.  Public LB: 0.04277\nModel2 -sorted by UserID then Date.  Public LB: 0.0425x\nModel3 -sorted by AdID then Date.  Public LB: 0.0425x\nThree models combined: 0.4235</p>\n\n<p>XGBoost was build over last 8 clicks of each UserID:\nModel 1, 2 and 3 were built using different features and hyperparameters, including all datasets, some time based features like counts per hour, history of clicks by User, by Ad. Also we created 6 xgb meta features using the dataset composed from clicks first to last 9 of each User to train. Also for features with too many levels like UserID and AdID we aplied a smoothing and also combined 2-way with other features to generate more featuers. Our XGB model 1 nd 2  have aprox. 50 feats and 3 about 100 feats.\nOur best XGB model 2 scored Public LB: 0.04108</p>\n\n<p>Golden Features: before SearchStream.tsv filter by ObjectType==3:   Sum of Objective==1 by SearchID, Sum of Objective==2 by SearchID, number of instances by SearchID.   Combine these 3 features.</p>\n\n<p>We trained and validated using the last click of each UserID. \nCombining XGB and FTRL scored 0.04088 public and private 0.04107.\nWorking with that dataset using a low end hardware is very hard :-( so we worked hard only in the last 7 days of the competition.</p>",
      "rawMarkdown": "Congratulations to the winners!!! That was a unique big data competition!\r\n\r\nOur aproach is based in two algos: FTRL and XGBoost. Basically we build 3 models  for each algo, then found some weights to combine then all.\r\n\r\nFTRL was built over all dataset:\r\nModel1 -sorted randomly.  Public LB: 0.04277\r\nModel2 -sorted by UserID then Date.  Public LB: 0.0425x\r\nModel3 -sorted by AdID then Date.  Public LB: 0.0425x\r\nThree models combined: 0.4235\r\n\r\nXGBoost was build over last 8 clicks of each UserID:\r\nModel 1, 2 and 3 were built using different features and hyperparameters, including all datasets, some time based features like counts per hour, history of clicks by User, by Ad. Also we created 6 xgb meta features using the dataset composed from clicks first to last 9 of each User to train. Also for features with too many levels like UserID and AdID we aplied a smoothing and also combined 2-way with other features to generate more featuers. Our XGB model 1 nd 2  have aprox. 50 feats and 3 about 100 feats.\r\nOur best XGB model 2 scored Public LB: 0.04108\r\n\r\nGolden Features: before SearchStream.tsv filter by ObjectType==3:   Sum of Objective==1 by SearchID, Sum of Objective==2 by SearchID, number of instances by SearchID.   Combine these 3 features.\r\n\r\nWe trained and validated using the last click of each UserID. \r\nCombining XGB and FTRL scored 0.04088 public and private 0.04107.\r\nWorking with that dataset using a low end hardware is very hard :-( so we worked hard only in the last 7 days of the competition.",
      "votes": null
    },
    {
      "id": "87434",
      "postDate": "07/29/2015 14:36:42",
      "content": "<p>Thank you all !  I will wait for your codes : D</p>",
      "rawMarkdown": "Thank you all !  I will wait for your codes : D",
      "votes": null
    },
    {
      "id": "87438",
      "postDate": "07/29/2015 15:17:44",
      "content": "<p>Congrats to the winners, and to everyone who participated for helping push the limits on this difficult dataset.  </p>\n\n<p>Even though I'd been using python for the last few comps, I ended up coming back to C++ for this one due to the size of the data and the memory restrictions on my laptop. </p>\n\n<p>Best submission was a simple average of a custom sparse-input NN using FTRL, with one 10 node hidden layer, and an FFM (using the --on-disk flag) on similar set of features. Both scored about the same on the public LB, and the blend only improved by 0.0001</p>\n\n<p>In terms of features, I ended up with 7 numeric, and 43 categorical (each hashed into a100k bucket), some of which were 2 and 3 way interactions. As others mentioned, the count of each type of Ad in the search instance was useful (premise being e.g. you might be more visually drawn to the highlighted ads when they are there), along with hash between them and their sum. Some numeric features included log(1+query_length), log(1+title_length), log(1+HistCTR*10000 ), number of times this user has seen this ad before, plus whether this user has already clicked on this ad before. \nAlso, since the error rate was so different between previously seen users, and new users, I also found it helpful to include previous search count by user and/or previous context ad impressions by user (hoping the algo will bias accordingly)</p>\n\n<p>A subset of the numeric features were also included as categorical features on instances when the count of the specific value exceeded 100)</p>\n\n<p>The most frustrating aspect for me was how slow it was to iterate and try new ideas, so a lot of my time was spent re-engineering parts of the code to make the iteration time faster (e.g. using custom bitfields for the records to save as many bits of memory as possible; sorting by UserID+Time prior to feature construction to minimize cache misses and disk fetches; parsing then persisting data in binary format post-processing at various stages to save time re-parsing; etc. if anyone is interested in more details, hit me up) </p>",
      "rawMarkdown": "Congrats to the winners, and to everyone who participated for helping push the limits on this difficult dataset.  \r\n\r\nEven though I'd been using python for the last few comps, I ended up coming back to C++ for this one due to the size of the data and the memory restrictions on my laptop. \r\n\r\nBest submission was a simple average of a custom sparse-input NN using FTRL, with one 10 node hidden layer, and an FFM (using the --on-disk flag) on similar set of features. Both scored about the same on the public LB, and the blend only improved by 0.0001\r\n\r\nIn terms of features, I ended up with 7 numeric, and 43 categorical (each hashed into a100k bucket), some of which were 2 and 3 way interactions. As others mentioned, the count of each type of Ad in the search instance was useful (premise being e.g. you might be more visually drawn to the highlighted ads when they are there), along with hash between them and their sum. Some numeric features included log(1+query_length), log(1+title_length), log(1+HistCTR*10000 ), number of times this user has seen this ad before, plus whether this user has already clicked on this ad before. \r\nAlso, since the error rate was so different between previously seen users, and new users, I also found it helpful to include previous search count by user and/or previous context ad impressions by user (hoping the algo will bias accordingly)\r\n\r\nA subset of the numeric features were also included as categorical features on instances when the count of the specific value exceeded 100)\r\n\r\nThe most frustrating aspect for me was how slow it was to iterate and try new ideas, so a lot of my time was spent re-engineering parts of the code to make the iteration time faster (e.g. using custom bitfields for the records to save as many bits of memory as possible; sorting by UserID+Time prior to feature construction to minimize cache misses and disk fetches; parsing then persisting data in binary format post-processing at various stages to save time re-parsing; etc. if anyone is interested in more details, hit me up)",
      "votes": null
    },
    {
      "id": "87451",
      "postDate": "07/29/2015 16:07:45",
      "content": "<p>[quote=rcarson;87358]</p>\n\n<p>I'm especially interested in the sampling and quick validation strategies. Thank you all in advance.</p>\n\n<p>[/quote]</p>\n\n<p>As long as you are using an online method, you can validate quickly on a small portion of the dataset by iterating directly on an sql cursor. Just save the ids to be validated to file, and then use progressive validation on part of the dataset that is close to the test dates.</p>\n\n<p>For our solution, we trained over a single pass of the data by using sql queries and updating a different model based on whether the search query was blank or not, and how many total searches the user had up until that point. These different user types had quite different click rate distributions, and since there was plenty of data you could train different portions according to different rules.</p>",
      "rawMarkdown": "[quote=rcarson;87358]\r\n\r\nI'm especially interested in the sampling and quick validation strategies. Thank you all in advance.\r\n\r\n[/quote]\r\n\r\nAs long as you are using an online method, you can validate quickly on a small portion of the dataset by iterating directly on an sql cursor. Just save the ids to be validated to file, and then use progressive validation on part of the dataset that is close to the test dates.\r\n\r\nFor our solution, we trained over a single pass of the data by using sql queries and updating a different model based on whether the search query was blank or not, and how many total searches the user had up until that point. These different user types had quite different click rate distributions, and since there was plenty of data you could train different portions according to different rules.",
      "votes": null
    },
    {
      "id": "87457",
      "postDate": "07/29/2015 16:19:10",
      "content": "<p>Congrats to the winners, the combat among top three player is really a close one!</p>\n\n<p>Reading the summary of some of the top ten players, it seems the <em>golden feature</em> is: </p>\n\n<p><strong>&quot;the count of each type of Ad in the search instance along with hash between them and their sum&quot;</strong></p>\n\n<p>this is actually a meaningful feature, sadly I never thought about generating new features from the original dataset once I filter out the objectiveType 1 and 2 to get a csv file. </p>\n\n<p>But after this competition, I finally know about  SQL, I even spend some time to read a book about SQL. </p>\n\n<p>Lessons learned from this competition: </p>\n\n<ol>\n<li>need to find a reliable validation method, I join this competition rather late, thus only used the public leaderboard as validation set to valid my FTRL models.</li>\n<li>when dataset is big, try find ways to samples it without damaging the performance much, so it will be easier to try out different ideas. Otherwise it is too painful and time consuming to try different algorithms and different feature engineering ideas.</li>\n</ol>",
      "rawMarkdown": "Congrats to the winners, the combat among top three player is really a close one!\r\n\r\nReading the summary of some of the top ten players, it seems the *golden feature* is: \r\n\r\n**\"the count of each type of Ad in the search instance along with hash between them and their sum\"**\r\n\r\nthis is actually a meaningful feature, sadly I never thought about generating new features from the original dataset once I filter out the objectiveType 1 and 2 to get a csv file. \r\n\r\nBut after this competition, I finally know about  SQL, I even spend some time to read a book about SQL. \r\n\r\nLessons learned from this competition: \r\n\r\n 1. need to find a reliable validation method, I join this competition rather late, thus only used the public leaderboard as validation set to valid my FTRL models.\r\n 2.  when dataset is big, try find ways to samples it without damaging the performance much, so it will be easier to try out different ideas. Otherwise it is too painful and time consuming to try different algorithms and different feature engineering ideas.",
      "votes": null
    },
    {
      "id": "87521",
      "postDate": "07/30/2015 00:31:57",
      "content": "<p>For those who used VW, did all of you shuffle the data in some way? I think my issue might have been I didn't shuffle the data at all and just used the original order it was presented. It might have also been an issue on FTRL. </p>",
      "rawMarkdown": "For those who used VW, did all of you shuffle the data in some way? I think my issue might have been I didn't shuffle the data at all and just used the original order it was presented. It might have also been an issue on FTRL.",
      "votes": null
    },
    {
      "id": "87522",
      "postDate": "07/30/2015 00:53:22",
      "content": "<p>@Mike Kim,\n Ordering the trainset by UserID,Date and AdID,Date improved my FTRL by 0.00050   :-)\nBut in my case I trained using mini batches FTRL, each User by time and each Ad by time</p>",
      "rawMarkdown": "Mike Kim,\r\n Ordering the trainset by UserID,Date and AdID,Date improved my FTRL by 0.00050   :-)\r\nBut in my case I trained using mini batches FTRL, each User by time and each Ad by time",
      "votes": null
    },
    {
      "id": "87523",
      "postDate": "07/30/2015 00:54:50",
      "content": "<p>Thanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?</p>",
      "rawMarkdown": "Thanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?",
      "votes": null
    },
    {
      "id": "87562",
      "postDate": "07/30/2015 08:31:43",
      "content": "<p>[quote=Mike Kim;87523]</p>\n\n<p>Thanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?</p>\n\n<p>[/quote]</p>\n\n<p>I used something similar in the Avazu contest, if you are interested, you can check my presentation <a href=\"http://mech.math.msu.su/~efimov/files/ClickThroughRate2015.pdf\">http://mech.math.msu.su/~efimov/files/ClickThroughRate2015.pdf</a> about the solution (there are some slides about Batch FTRL). The principle is to sort combined dataset (train and test) by some fields and apply the FTRL algorithm. Then it will work in the following way: the algorithm is trained on the first batch from the train set and make prediction on the first batch from the test set. Then it starts training on the second batch from the train set starting from coefficients from the first batch and so on.</p>",
      "rawMarkdown": "[quote=Mike Kim;87523]\r\n\r\nThanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?\r\n\r\n[/quote]\r\n\r\nI used something similar in the Avazu contest, if you are interested, you can check my presentation http://mech.math.msu.su/~efimov/files/ClickThroughRate2015.pdf about the solution (there are some slides about Batch FTRL). The principle is to sort combined dataset (train and test) by some fields and apply the FTRL algorithm. Then it will work in the following way: the algorithm is trained on the first batch from the train set and make prediction on the first batch from the test set. Then it starts training on the second batch from the train set starting from coefficients from the first batch and so on.",
      "votes": null
    },
    {
      "id": "87584",
      "postDate": "07/30/2015 12:50:42",
      "content": "<p>@Mike: Dmitry explained all!\nMy FTRL is based in the original Tingrtu code and executed via pypy.\n<a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory\">https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory</a></p>",
      "rawMarkdown": "Mike: Dmitry explained all!\r\nMy FTRL is based in the original Tingrtu code and executed via pypy.\r\nhttps://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory",
      "votes": null
    },
    {
      "id": "88055",
      "postDate": "08/03/2015 05:47:00",
      "content": "<p>I learned so much from this competition and thanks everyone for all the amazing ideas sharing!! </p>\n\n<p>I still have one question I would like to ask here, in practice, will a 0.0402 logloss model be considered as a &quot;good&quot; or &quot;useful&quot; model? And how big a difference it actually is between, say, 0.041 and 0.043?  Thanks!</p>",
      "rawMarkdown": "I learned so much from this competition and thanks everyone for all the amazing ideas sharing!! \r\n\r\nI still have one question I would like to ask here, in practice, will a 0.0402 logloss model be considered as a \"good\" or \"useful\" model? And how big a difference it actually is between, say, 0.041 and 0.043?  Thanks!",
      "votes": null
    },
    {
      "id": "89533",
      "postDate": "08/16/2015 22:09:48",
      "content": "<p>Sorry for the delay, our model is uploaded now: <a href=\"https://github.com/diefimov/avito_context_click_2015\">https://github.com/diefimov/avito_context_click_2015</a></p>",
      "rawMarkdown": "Sorry for the delay, our model is uploaded now: https://github.com/diefimov/avito_context_click_2015",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 87362,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "07/29/2015 00:45:33",
      "content": "<p>Congratulations to the top 10 and all other competitors that made this exciting.  We tried pushing for top 10 but couldn't make it at the end.  In fact, our single Vowpal Wabbit model would have still placed us in 11th still.  2 important things we did in this competition:</p>\n\n<ol>\n<li>Sorted users' searches by time so that each place (e.g., last search by user, second to last search by user, etc) were stored in separate training files.  Then trained on the first 9 of the last 10 searches in training (in the order) and validating on the last search before test.  This reduced the training to only ~30 Million records.  This correlated near perfectly with the Public LB.  Final submission was trained on only the last 99 searches in the training set (~112M records).</li>\n<li>Number of ads, and number of types of ads (Types 1, 2, &amp; 3) on the search provided one of the biggest boosts.</li>\n</ol>\n\n<p>We should have used the VisitStreams and PhoneRequestStream, but ran out of time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87372,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "07/29/2015 01:49:59",
      "content": "<p>I couldn't figure this competition out either. I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? </p>\n\n<p>I did, however, use all the data including both the phone and visits stream. I also engineered features including bag of Russian words (attached). I'm not sure if it was a correct way of going about it, but it seemed to give a small boost.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87373,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "07/29/2015 02:10:21",
      "content": "<p>[quote=Mike Kim;87372]</p>\n\n<p>I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? </p>\n\n<p>[/quote]\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:</p>\n\n<pre><code>vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87374,
      "author_name": "gzsiceberg",
      "author_url": "",
      "post_date": "07/29/2015 02:21:19",
      "content": "<p>Congratulations to the winners,  I'm very lucky in the private leaderboard, I spent a lot of time pushing myself in top2, but when I thought a idea to improve my scores, I always saw that Owen and Dmitry &amp; Leustagos improve their scores again, this is a really hard competition.</p>\n\n<p>I list some key points.</p>\n\n<p>(1). negative sampling(0.1) to reduce data size(about 20M records).</p>\n\n<p>(2). a lot of features, like number of ads, rank by price for every ads, position type(all ads' position in one tuple, then use hash value as a category feature), show count and click count for every users,  and so on. But I also don't use VisitStreams and PhoneRequestStream, maybe I should try it.\nBy the way, I just test new features in first 10M records of trainSearchStream.tsv, it looks like very effective.</p>\n\n<p>(3). ensemble. FFM, FM, GBDT as base models, then use NN and GBDT to combine it. Finally, I average my best 3 submissions. </p>\n\n<p>best model is a FFM model, it give me 0.04061 in public leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87377,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "07/29/2015 02:44:17",
      "content": "<p>I don't think that's the issue since I hashed tricked every single string in Python when parsing the file for VW format. Then again my parsing script could have been buggy. I kept getting saturated predictions of exactly 0 or 1 or (-50, or 50 when raw?) or a seg fault within VW. I think it's a system problem since I recently got a new computer (64GB ram which seems like not enough) and how to reinstall VW and everything else.\n[quote=David Shinn;87373]</p>\n\n<p>[quote=Mike Kim;87372]</p>\n\n<p>I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? </p>\n\n<p>[/quote]\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:</p>\n\n<pre><code>vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0\n</code></pre>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87379,
      "author_name": "skirpichenko",
      "author_url": "",
      "post_date": "07/29/2015 03:14:46",
      "content": "<p>Congratulations to the winners!</p>\n\n<p>I spent at least half of the time trying to employ VisitsStream and RequestStream in my model (FFM).   I tried a number of straightforward features like AD words, stemmed words, category, time etc. But each time I was getting only marginal improvement. The last and most promising idea was to compose  'sentences' from the titles of ADs consequently viewed by a user and train word2vec on it. Then it can be used in many ways to generate features based on distance between ADs or query to AD or ads classification etc.  Unfortunately I ran out of time implementing this but preliminary tests on sample showed impressive improvement.</p>\n\n<p>If somebody successfully used personalization data I would be happy if you share a way you did it. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87405,
      "author_name": "morandi",
      "author_url": "",
      "post_date": "07/29/2015 11:33:31",
      "content": "<p>Congratulation to winners and all top teams! Well done!</p>\n\n<p>This was a very interesting competition! Improvements on validation set reflected the same improvements on LB. Our validation set was the same as that described here: <a href=\"https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15367/proper-validation-set/86077#post86077\">https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15367/proper-validation-set/86077#post86077</a>.\nAnd it works perfectly:)</p>\n\n<p>Our approach:</p>\n\n<p>1) My single ftrl model with basic features: AdID, UserID, IPID, Position, Price, &#8230; + the second- and the third-order interactions between some of this features + scores similarity between SearchQuery and Title and SearchParams and Params. It gets Public LB score below 0.044.</p>\n\n<p>2) Adding to this model only one new feature PositionFactor from my teammate Alexander. It gave the incredible increments and the model scored  ~ 0.0418 on Public LB. </p>\n\n<p>PositionFactor = hash( [Position 1_place:ObjectType 2_place:ObjectType 6_place:ObjectType 7_place:ObjectType 8_place:ObjectType]) , where \nPosition is the position of the given context ad in search result page.\nk_place:ObjectType is the position and type others ads in search result page (if there are data about them).</p>\n\n<p>3) Alexander's models vw+xgb and vw+rf give him ~ 0.0424 on Public LB</p>\n\n<p>4) After tuning hyperparameters and finding the weights for linear combination of our solutions we achieved 0.04137 on Public LB. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87421,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/29/2015 13:09:57",
      "content": "<p>Congratulations to the winners!!! That was a unique big data competition!</p>\n\n<p>Our aproach is based in two algos: FTRL and XGBoost. Basically we build 3 models  for each algo, then found some weights to combine then all.</p>\n\n<p>FTRL was built over all dataset:\nModel1 -sorted randomly.  Public LB: 0.04277\nModel2 -sorted by UserID then Date.  Public LB: 0.0425x\nModel3 -sorted by AdID then Date.  Public LB: 0.0425x\nThree models combined: 0.4235</p>\n\n<p>XGBoost was build over last 8 clicks of each UserID:\nModel 1, 2 and 3 were built using different features and hyperparameters, including all datasets, some time based features like counts per hour, history of clicks by User, by Ad. Also we created 6 xgb meta features using the dataset composed from clicks first to last 9 of each User to train. Also for features with too many levels like UserID and AdID we aplied a smoothing and also combined 2-way with other features to generate more featuers. Our XGB model 1 nd 2  have aprox. 50 feats and 3 about 100 feats.\nOur best XGB model 2 scored Public LB: 0.04108</p>\n\n<p>Golden Features: before SearchStream.tsv filter by ObjectType==3:   Sum of Objective==1 by SearchID, Sum of Objective==2 by SearchID, number of instances by SearchID.   Combine these 3 features.</p>\n\n<p>We trained and validated using the last click of each UserID. \nCombining XGB and FTRL scored 0.04088 public and private 0.04107.\nWorking with that dataset using a low end hardware is very hard :-( so we worked hard only in the last 7 days of the competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87434,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "07/29/2015 14:36:42",
      "content": "<p>Thank you all !  I will wait for your codes : D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87438,
      "author_name": "ashhafez",
      "author_url": "",
      "post_date": "07/29/2015 15:17:44",
      "content": "<p>Congrats to the winners, and to everyone who participated for helping push the limits on this difficult dataset.  </p>\n\n<p>Even though I'd been using python for the last few comps, I ended up coming back to C++ for this one due to the size of the data and the memory restrictions on my laptop. </p>\n\n<p>Best submission was a simple average of a custom sparse-input NN using FTRL, with one 10 node hidden layer, and an FFM (using the --on-disk flag) on similar set of features. Both scored about the same on the public LB, and the blend only improved by 0.0001</p>\n\n<p>In terms of features, I ended up with 7 numeric, and 43 categorical (each hashed into a100k bucket), some of which were 2 and 3 way interactions. As others mentioned, the count of each type of Ad in the search instance was useful (premise being e.g. you might be more visually drawn to the highlighted ads when they are there), along with hash between them and their sum. Some numeric features included log(1+query_length), log(1+title_length), log(1+HistCTR*10000 ), number of times this user has seen this ad before, plus whether this user has already clicked on this ad before. \nAlso, since the error rate was so different between previously seen users, and new users, I also found it helpful to include previous search count by user and/or previous context ad impressions by user (hoping the algo will bias accordingly)</p>\n\n<p>A subset of the numeric features were also included as categorical features on instances when the count of the specific value exceeded 100)</p>\n\n<p>The most frustrating aspect for me was how slow it was to iterate and try new ideas, so a lot of my time was spent re-engineering parts of the code to make the iteration time faster (e.g. using custom bitfields for the records to save as many bits of memory as possible; sorting by UserID+Time prior to feature construction to minimize cache misses and disk fetches; parsing then persisting data in binary format post-processing at various stages to save time re-parsing; etc. if anyone is interested in more details, hit me up) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87451,
      "author_name": "dkaylor",
      "author_url": "",
      "post_date": "07/29/2015 16:07:45",
      "content": "<p>[quote=rcarson;87358]</p>\n\n<p>I'm especially interested in the sampling and quick validation strategies. Thank you all in advance.</p>\n\n<p>[/quote]</p>\n\n<p>As long as you are using an online method, you can validate quickly on a small portion of the dataset by iterating directly on an sql cursor. Just save the ids to be validated to file, and then use progressive validation on part of the dataset that is close to the test dates.</p>\n\n<p>For our solution, we trained over a single pass of the data by using sql queries and updating a different model based on whether the search query was blank or not, and how many total searches the user had up until that point. These different user types had quite different click rate distributions, and since there was plenty of data you could train different portions according to different rules.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87457,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "07/29/2015 16:19:10",
      "content": "<p>Congrats to the winners, the combat among top three player is really a close one!</p>\n\n<p>Reading the summary of some of the top ten players, it seems the <em>golden feature</em> is: </p>\n\n<p><strong>&quot;the count of each type of Ad in the search instance along with hash between them and their sum&quot;</strong></p>\n\n<p>this is actually a meaningful feature, sadly I never thought about generating new features from the original dataset once I filter out the objectiveType 1 and 2 to get a csv file. </p>\n\n<p>But after this competition, I finally know about  SQL, I even spend some time to read a book about SQL. </p>\n\n<p>Lessons learned from this competition: </p>\n\n<ol>\n<li>need to find a reliable validation method, I join this competition rather late, thus only used the public leaderboard as validation set to valid my FTRL models.</li>\n<li>when dataset is big, try find ways to samples it without damaging the performance much, so it will be easier to try out different ideas. Otherwise it is too painful and time consuming to try different algorithms and different feature engineering ideas.</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87521,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "07/30/2015 00:31:57",
      "content": "<p>For those who used VW, did all of you shuffle the data in some way? I think my issue might have been I didn't shuffle the data at all and just used the original order it was presented. It might have also been an issue on FTRL. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87522,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/30/2015 00:53:22",
      "content": "<p>@Mike Kim,\n Ordering the trainset by UserID,Date and AdID,Date improved my FTRL by 0.00050   :-)\nBut in my case I trained using mini batches FTRL, each User by time and each Ad by time</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87523,
      "author_name": "mikeskim",
      "author_url": "",
      "post_date": "07/30/2015 00:54:50",
      "content": "<p>Thanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87562,
      "author_name": "efimov",
      "author_url": "",
      "post_date": "07/30/2015 08:31:43",
      "content": "<p>[quote=Mike Kim;87523]</p>\n\n<p>Thanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?</p>\n\n<p>[/quote]</p>\n\n<p>I used something similar in the Avazu contest, if you are interested, you can check my presentation <a href=\"http://mech.math.msu.su/~efimov/files/ClickThroughRate2015.pdf\">http://mech.math.msu.su/~efimov/files/ClickThroughRate2015.pdf</a> about the solution (there are some slides about Batch FTRL). The principle is to sort combined dataset (train and test) by some fields and apply the FTRL algorithm. Then it will work in the following way: the algorithm is trained on the first batch from the train set and make prediction on the first batch from the test set. Then it starts training on the second batch from the train set starting from coefficients from the first batch and so on.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87584,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/30/2015 12:50:42",
      "content": "<p>@Mike: Dmitry explained all!\nMy FTRL is based in the original Tingrtu code and executed via pypy.\n<a href=\"https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory\">https://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88055,
      "author_name": "hitchpy",
      "author_url": "",
      "post_date": "08/03/2015 05:47:00",
      "content": "<p>I learned so much from this competition and thanks everyone for all the amazing ideas sharing!! </p>\n\n<p>I still have one question I would like to ask here, in practice, will a 0.0402 logloss model be considered as a &quot;good&quot; or &quot;useful&quot; model? And how big a difference it actually is between, say, 0.041 and 0.043?  Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 89533,
      "author_name": "efimov",
      "author_url": "",
      "post_date": "08/16/2015 22:09:48",
      "content": "<p>Sorry for the delay, our model is uploaded now: <a href=\"https://github.com/diefimov/avito_context_click_2015\">https://github.com/diefimov/avito_context_click_2015</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "87358": "Great contest! Congratulations to top players! To be honest I studied several previous winning solutions but still cannot figure out this one. I'm especially interested in the sampling and quick validation strategies. Thank you all in advance.",
    "87362": "Congratulations to the top 10 and all other competitors that made this exciting.  We tried pushing for top 10 but couldn't make it at the end.  In fact, our single Vowpal Wabbit model would have still placed us in 11th still.  2 important things we did in this competition:\r\n\r\n1. Sorted users' searches by time so that each place (e.g., last search by user, second to last search by user, etc) were stored in separate training files.  Then trained on the first 9 of the last 10 searches in training (in the order) and validating on the last search before test.  This reduced the training to only ~30 Million records.  This correlated near perfectly with the Public LB.  Final submission was trained on only the last 99 searches in the training set (~112M records).\r\n2. Number of ads, and number of types of ads (Types 1, 2, & 3) on the search provided one of the biggest boosts.\r\n\r\nWe should have used the VisitStreams and PhoneRequestStream, but ran out of time.",
    "87372": "I couldn't figure this competition out either. I couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? \r\n\r\nI did, however, use all the data including both the phone and visits stream. I also engineered features including bag of Russian words (attached). I'm not sure if it was a correct way of going about it, but it seemed to give a small boost.",
    "87373": "[quote=Mike Kim;87372]\r\n\r\nI couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? \r\n\r\n[/quote]\r\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:\r\n\r\n    vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0",
    "87374": "Congratulations to the winners,  I'm very lucky in the private leaderboard, I spent a lot of time pushing myself in top2, but when I thought a idea to improve my scores, I always saw that Owen and Dmitry & Leustagos improve their scores again, this is a really hard competition.\r\n\r\nI list some key points.\r\n\r\n(1). negative sampling(0.1) to reduce data size(about 20M records).\r\n\r\n(2). a lot of features, like number of ads, rank by price for every ads, position type(all ads' position in one tuple, then use hash value as a category feature), show count and click count for every users,  and so on. But I also don't use VisitStreams and PhoneRequestStream, maybe I should try it.\r\nBy the way, I just test new features in first 10M records of trainSearchStream.tsv, it looks like very effective.\r\n\r\n(3). ensemble. FFM, FM, GBDT as base models, then use NN and GBDT to combine it. Finally, I average my best 3 submissions. \r\n\r\nbest model is a FFM model, it give me 0.04061 in public leaderboard.",
    "87377": "I don't think that's the issue since I hashed tricked every single string in Python when parsing the file for VW format. Then again my parsing script could have been buggy. I kept getting saturated predictions of exactly 0 or 1 or (-50, or 50 when raw?) or a seg fault within VW. I think it's a system problem since I recently got a new computer (64GB ram which seems like not enough) and how to reinstall VW and everything else.\r\n[quote=David Shinn;87373]\r\n\r\n[quote=Mike Kim;87372]\r\n\r\nI couldn't even get Vowpal to work (and I have used it in the past many times) at all. Was there some trick to it? \r\n\r\n[/quote]\r\n@Mike, have you considered the text encoding issues.  vw wouldn't work if I weren't careful.  In Python2.7, when we read the data, we had to decode('utf-8') and then when writing to the vw file format, encode('utf-8').  We were running vw full tilt on -b 30 (on macbook pro with 16 GB RAM), but that wasn't enough because on our final vw submission, it trains on 13,781,306,252 features so there were certainly a lot of hash collisions.  We were pretty much using one feature per namespace and had namespaces from A-Z, a-z, 0-9, !@#$%.  This is what the training command line looked like with all of the switches:\r\n\r\n    vw --loss_function logistic -b 30 --power_t 0.4 -l 0.6 -k -q ak -q FG -q FM -q aM -q at -q aF -q ac -q Fh -q tx -q tG -q Mt -q tq -q IJ --cubic aFk --cubic aFM -q 4k -q 43 -q 42 -q 4a -q 4M --redefine 0:=YCLEQSITU --ignore 0\r\n\r\n[/quote]",
    "87379": "Congratulations to the winners!\r\n\r\nI spent at least half of the time trying to employ VisitsStream and RequestStream in my model (FFM).   I tried a number of straightforward features like AD words, stemmed words, category, time etc. But each time I was getting only marginal improvement. The last and most promising idea was to compose  'sentences' from the titles of ADs consequently viewed by a user and train word2vec on it. Then it can be used in many ways to generate features based on distance between ADs or query to AD or ads classification etc.  Unfortunately I ran out of time implementing this but preliminary tests on sample showed impressive improvement.\r\n\r\nIf somebody successfully used personalization data I would be happy if you share a way you did it.",
    "87405": "Congratulation to winners and all top teams! Well done!\r\n\r\nThis was a very interesting competition! Improvements on validation set reflected the same improvements on LB. Our validation set was the same as that described here: https://www.kaggle.com/c/avito-context-ad-clicks/forums/t/15367/proper-validation-set/86077#post86077.\r\nAnd it works perfectly:)\r\n\r\nOur approach:\r\n\r\n1) My single ftrl model with basic features: AdID, UserID, IPID, Position, Price, … + the second- and the third-order interactions between some of this features + scores similarity between SearchQuery and Title and SearchParams and Params. It gets Public LB score below 0.044.\r\n\r\n2) Adding to this model only one new feature PositionFactor from my teammate Alexander. It gave the incredible increments and the model scored  ~ 0.0418 on Public LB. \r\n\r\nPositionFactor = hash( [Position 1_place:ObjectType 2_place:ObjectType 6_place:ObjectType 7_place:ObjectType 8_place:ObjectType]) , where \r\nPosition is the position of the given context ad in search result page.\r\nk_place:ObjectType is the position and type others ads in search result page (if there are data about them).\r\n\r\n3) Alexander's models vw+xgb and vw+rf give him ~ 0.0424 on Public LB\r\n\r\n4) After tuning hyperparameters and finding the weights for linear combination of our solutions we achieved 0.04137 on Public LB.",
    "87421": "Congratulations to the winners!!! That was a unique big data competition!\r\n\r\nOur aproach is based in two algos: FTRL and XGBoost. Basically we build 3 models  for each algo, then found some weights to combine then all.\r\n\r\nFTRL was built over all dataset:\r\nModel1 -sorted randomly.  Public LB: 0.04277\r\nModel2 -sorted by UserID then Date.  Public LB: 0.0425x\r\nModel3 -sorted by AdID then Date.  Public LB: 0.0425x\r\nThree models combined: 0.4235\r\n\r\nXGBoost was build over last 8 clicks of each UserID:\r\nModel 1, 2 and 3 were built using different features and hyperparameters, including all datasets, some time based features like counts per hour, history of clicks by User, by Ad. Also we created 6 xgb meta features using the dataset composed from clicks first to last 9 of each User to train. Also for features with too many levels like UserID and AdID we aplied a smoothing and also combined 2-way with other features to generate more featuers. Our XGB model 1 nd 2  have aprox. 50 feats and 3 about 100 feats.\r\nOur best XGB model 2 scored Public LB: 0.04108\r\n\r\nGolden Features: before SearchStream.tsv filter by ObjectType==3:   Sum of Objective==1 by SearchID, Sum of Objective==2 by SearchID, number of instances by SearchID.   Combine these 3 features.\r\n\r\nWe trained and validated using the last click of each UserID. \r\nCombining XGB and FTRL scored 0.04088 public and private 0.04107.\r\nWorking with that dataset using a low end hardware is very hard :-( so we worked hard only in the last 7 days of the competition.",
    "87434": "Thank you all !  I will wait for your codes : D",
    "87438": "Congrats to the winners, and to everyone who participated for helping push the limits on this difficult dataset.  \r\n\r\nEven though I'd been using python for the last few comps, I ended up coming back to C++ for this one due to the size of the data and the memory restrictions on my laptop. \r\n\r\nBest submission was a simple average of a custom sparse-input NN using FTRL, with one 10 node hidden layer, and an FFM (using the --on-disk flag) on similar set of features. Both scored about the same on the public LB, and the blend only improved by 0.0001\r\n\r\nIn terms of features, I ended up with 7 numeric, and 43 categorical (each hashed into a100k bucket), some of which were 2 and 3 way interactions. As others mentioned, the count of each type of Ad in the search instance was useful (premise being e.g. you might be more visually drawn to the highlighted ads when they are there), along with hash between them and their sum. Some numeric features included log(1+query_length), log(1+title_length), log(1+HistCTR*10000 ), number of times this user has seen this ad before, plus whether this user has already clicked on this ad before. \r\nAlso, since the error rate was so different between previously seen users, and new users, I also found it helpful to include previous search count by user and/or previous context ad impressions by user (hoping the algo will bias accordingly)\r\n\r\nA subset of the numeric features were also included as categorical features on instances when the count of the specific value exceeded 100)\r\n\r\nThe most frustrating aspect for me was how slow it was to iterate and try new ideas, so a lot of my time was spent re-engineering parts of the code to make the iteration time faster (e.g. using custom bitfields for the records to save as many bits of memory as possible; sorting by UserID+Time prior to feature construction to minimize cache misses and disk fetches; parsing then persisting data in binary format post-processing at various stages to save time re-parsing; etc. if anyone is interested in more details, hit me up)",
    "87451": "[quote=rcarson;87358]\r\n\r\nI'm especially interested in the sampling and quick validation strategies. Thank you all in advance.\r\n\r\n[/quote]\r\n\r\nAs long as you are using an online method, you can validate quickly on a small portion of the dataset by iterating directly on an sql cursor. Just save the ids to be validated to file, and then use progressive validation on part of the dataset that is close to the test dates.\r\n\r\nFor our solution, we trained over a single pass of the data by using sql queries and updating a different model based on whether the search query was blank or not, and how many total searches the user had up until that point. These different user types had quite different click rate distributions, and since there was plenty of data you could train different portions according to different rules.",
    "87457": "Congrats to the winners, the combat among top three player is really a close one!\r\n\r\nReading the summary of some of the top ten players, it seems the *golden feature* is: \r\n\r\n**\"the count of each type of Ad in the search instance along with hash between them and their sum\"**\r\n\r\nthis is actually a meaningful feature, sadly I never thought about generating new features from the original dataset once I filter out the objectiveType 1 and 2 to get a csv file. \r\n\r\nBut after this competition, I finally know about  SQL, I even spend some time to read a book about SQL. \r\n\r\nLessons learned from this competition: \r\n\r\n 1. need to find a reliable validation method, I join this competition rather late, thus only used the public leaderboard as validation set to valid my FTRL models.\r\n 2.  when dataset is big, try find ways to samples it without damaging the performance much, so it will be easier to try out different ideas. Otherwise it is too painful and time consuming to try different algorithms and different feature engineering ideas.",
    "87521": "For those who used VW, did all of you shuffle the data in some way? I think my issue might have been I didn't shuffle the data at all and just used the original order it was presented. It might have also been an issue on FTRL.",
    "87522": "Mike Kim,\r\n Ordering the trainset by UserID,Date and AdID,Date improved my FTRL by 0.00050   :-)\r\nBut in my case I trained using mini batches FTRL, each User by time and each Ad by time",
    "87523": "Thanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?",
    "87562": "[quote=Mike Kim;87523]\r\n\r\nThanks. Can you explain exactly what a mini batch is in FTRL? I'm not familiar with this. Was your FTRL based upon VW or something else?\r\n\r\n[/quote]\r\n\r\nI used something similar in the Avazu contest, if you are interested, you can check my presentation http://mech.math.msu.su/~efimov/files/ClickThroughRate2015.pdf about the solution (there are some slides about Batch FTRL). The principle is to sort combined dataset (train and test) by some fields and apply the FTRL algorithm. Then it will work in the following way: the algorithm is trained on the first batch from the train set and make prediction on the first batch from the test set. Then it starts training on the second batch from the train set starting from coefficients from the first batch and so on.",
    "87584": "Mike: Dmitry explained all!\r\nMy FTRL is based in the original Tingrtu code and executed via pypy.\r\nhttps://www.kaggle.com/c/criteo-display-ad-challenge/forums/t/10322/beat-the-benchmark-with-less-then-200mb-of-memory",
    "88055": "I learned so much from this competition and thanks everyone for all the amazing ideas sharing!! \r\n\r\nI still have one question I would like to ask here, in practice, will a 0.0402 logloss model be considered as a \"good\" or \"useful\" model? And how big a difference it actually is between, say, 0.041 and 0.043?  Thanks!",
    "89533": "Sorry for the delay, our model is uploaded now: https://github.com/diefimov/avito_context_click_2015"
  },
  "source": "meta"
}