{
  "id": 15347,
  "title": "Best single model score?",
  "url": "/competitions/avito-context-ad-clicks/discussion/15347",
  "author_name": "",
  "post_date": "2015-07-18T20:36:29.497Z",
  "votes": null,
  "comment_count": 13,
  "views": 1681,
  "content": "<p>What's your best single model score? We could push ourselves a little more if the leaders throw some light on their score :-)</p>",
  "messages": [
    {
      "id": "85977",
      "postDate": "07/18/2015 20:36:29",
      "content": "<p>What's your best single model score? We could push ourselves a little more if the leaders throw some light on their score :-)</p>",
      "rawMarkdown": "What's your best single model score? We could push ourselves a little more if the leaders throw some light on their score :-)",
      "votes": null
    },
    {
      "id": "85979",
      "postDate": "07/18/2015 20:44:40",
      "content": "<p>my current LB score (0.04546) is a single ftrl model with basic features from the most of the 8 data tables provided.</p>\n\n<p>What is yours?</p>",
      "rawMarkdown": "my current LB score (0.04546) is a single ftrl model with basic features from the most of the 8 data tables provided.\r\n\r\nWhat is yours?",
      "votes": null
    },
    {
      "id": "85981",
      "postDate": "07/18/2015 21:00:52",
      "content": "<p>Same here. 0.04562 - single FTRL. With raw features fro most of the tables.</p>\n\n<p>I am trying feature selection and interactions but the iteration time is really high. Takes almost 55 minutes for 1 epoch on 190M records.</p>\n\n<p>Any suggestions on how to approach interactions?</p>",
      "rawMarkdown": "Same here. 0.04562 - single FTRL. With raw features fro most of the tables.\r\n\r\nI am trying feature selection and interactions but the iteration time is really high. Takes almost 55 minutes for 1 epoch on 190M records.\r\n\r\nAny suggestions on how to approach interactions?",
      "votes": null
    },
    {
      "id": "85985",
      "postDate": "07/18/2015 22:57:34",
      "content": "<blockquote>\n  <p>Any suggestions on how to approach interactions?</p>\n</blockquote>\n\n<p>Not competing yet, but try <a href=\"http://blogs.technet.com/b/machinelearning/archive/2014/09/24/online-learning-and-sub-linear-debugging.aspx\">sub-linear debugging</a>. </p>\n\n<p>Basically:</p>\n\n<ul>\n<li>try a run with a few interactions and see if the loss improves over 500k records. \n\n<blockquote>\n  <p>If so: cancel run, keep interactions and repeat, else: cancel run, discard interactions and repeat. </p>\n</blockquote></li>\n</ul>\n\n<p>Once you are done, make another run with all the interactions you found and confirm that it actually lowers progressive validation loss on complete train set. I think it may be considered a form of greedy forward feature interaction selection. </p>",
      "rawMarkdown": ">Any suggestions on how to approach interactions?\r\n\r\nNot competing yet, but try [sub-linear debugging][1]. \r\n\r\nBasically:\r\n\r\n*  try a run with a few interactions and see if the loss improves over 500k records. \r\n>If so: cancel run, keep interactions and repeat, else: cancel run, discard interactions and repeat. \r\n\r\nOnce you are done, make another run with all the interactions you found and confirm that it actually lowers progressive validation loss on complete train set. I think it may be considered a form of greedy forward feature interaction selection. \r\n\r\n\r\n  [1]: http://blogs.technet.com/b/machinelearning/archive/2014/09/24/online-learning-and-sub-linear-debugging.aspx",
      "votes": null
    },
    {
      "id": "85999",
      "postDate": "07/19/2015 01:49:33",
      "content": "<p>Its possible to get below 0.04100 with a single model...</p>",
      "rawMarkdown": "Its possible to get below 0.04100 with a single model...",
      "votes": null
    },
    {
      "id": "86003",
      "postDate": "07/19/2015 02:23:23",
      "content": "<p>Below 0.043 with a single FTRL.</p>",
      "rawMarkdown": "Below 0.043 with a single FTRL.",
      "votes": null
    },
    {
      "id": "86004",
      "postDate": "07/19/2015 02:27:39",
      "content": "<p>I must miss something. LB tells me ads masters found some pretty straightforward way to get under 0.44. I hope I could figure it out before the deadline. This is the most fun part of kaggle after all :D</p>",
      "rawMarkdown": "I must miss something. LB tells me ads masters found some pretty straightforward way to get under 0.44. I hope I could figure it out before the deadline. This is the most fun part of kaggle after all :D",
      "votes": null
    },
    {
      "id": "86007",
      "postDate": "07/19/2015 02:34:32",
      "content": "<p>Why I can only get 0.050 with the raw features from 6 tables(no SisitsStream and PhoneRequestsStream)\nand the simple FTRL from Abhishek's code(without any changes).\nI am wondering if I need to do some parameter tunning or maybe my features are messed up.</p>",
      "rawMarkdown": "Why I can only get 0.050 with the raw features from 6 tables(no SisitsStream and PhoneRequestsStream)\r\nand the simple FTRL from Abhishek's code(without any changes).\r\nI am wondering if I need to do some parameter tunning or maybe my features are messed up.",
      "votes": null
    },
    {
      "id": "86009",
      "postDate": "07/19/2015 03:01:28",
      "content": "<p>@SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P</p>",
      "rawMarkdown": "SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P",
      "votes": null
    },
    {
      "id": "86014",
      "postDate": "07/19/2015 03:18:01",
      "content": "<p>[quote=Leustagos;85999]</p>\n\n<p>Its possible to get below 0.04100 with a single model...</p>\n\n<p>[/quote]</p>\n\n<p>..., unfortunately Lucas isn't telling us what KIND of single model it is. </p>\n\n<p>:-)</p>",
      "rawMarkdown": "[quote=Leustagos;85999]\r\n\r\nIts possible to get below 0.04100 with a single model...\r\n\r\n[/quote]\r\n\r\n..., unfortunately Lucas isn't telling us what KIND of single model it is. \r\n\r\n\r\n:-)",
      "votes": null
    },
    {
      "id": "86025",
      "postDate": "07/19/2015 05:17:06",
      "content": "<p>I'm still struggling getting a working validation set, which most would interpret as the set that produces similar scores to the LB. But after thinking about this, I think it's wrong to measure the quality of the validation set against the LB score. The reason is that if the scores are equal, it only guarantees that the validation set has a similar distribution of 1's versus 0's?</p>\n\n<p>With the LB score based on 30% of the data, if the current 30% is <em>not</em> representative of the entire LB set, then the LB position doesn't necessarily represent the winner. Also, it probably means that the best model/parameter set isn't necessarily a winning one.</p>\n\n<p>Are these observations correct and can this competition for that reason gives us some very interesting surprises at the end?</p>",
      "rawMarkdown": "I'm still struggling getting a working validation set, which most would interpret as the set that produces similar scores to the LB. But after thinking about this, I think it's wrong to measure the quality of the validation set against the LB score. The reason is that if the scores are equal, it only guarantees that the validation set has a similar distribution of 1's versus 0's?\r\n\r\nWith the LB score based on 30% of the data, if the current 30% is *not* representative of the entire LB set, then the LB position doesn't necessarily represent the winner. Also, it probably means that the best model/parameter set isn't necessarily a winning one.\r\n\r\nAre these observations correct and can this competition for that reason gives us some very interesting surprises at the end?",
      "votes": null
    },
    {
      "id": "86026",
      "postDate": "07/19/2015 05:30:05",
      "content": "<p>[quote=rcarson;86009]</p>\n\n<p>@SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P</p>\n\n<p>[/quote]\nThanks, rcarson:)\nI really should be more careful...</p>",
      "rawMarkdown": "[quote=rcarson;86009]\r\n\r\n@SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P\r\n\r\n[/quote]\r\nThanks, rcarson:)\r\nI really should be more careful...",
      "votes": null
    },
    {
      "id": "86040",
      "postDate": "07/19/2015 08:47:03",
      "content": "<p>[quote=Owen;86014]</p>\n\n<p>[quote=Leustagos;85999]</p>\n\n<p>Its possible to get below 0.04100 with a single model...</p>\n\n<p>[/quote]</p>\n\n<p>..., unfortunately Lucas isn't telling us what KIND of single model it is. </p>\n\n<p>:-)</p>\n\n<p>[/quote]</p>\n\n<p>Haha! I think, going by the # of submissions of D&amp;L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)</p>",
      "rawMarkdown": "[quote=Owen;86014]\r\n\r\n[quote=Leustagos;85999]\r\n\r\nIts possible to get below 0.04100 with a single model...\r\n\r\n[/quote]\r\n\r\n..., unfortunately Lucas isn't telling us what KIND of single model it is. \r\n\r\n\r\n:-)\r\n\r\n[/quote]\r\n\r\nHaha! I think, going by the # of submissions of D&L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)",
      "votes": null
    },
    {
      "id": "86069",
      "postDate": "07/19/2015 14:05:52",
      "content": "<p>[quote=binga;86040]</p>\n\n<p>[quote=Owen;86014]</p>\n\n<p>[quote=Leustagos;85999]</p>\n\n<p>Its possible to get below 0.04100 with a single model...</p>\n\n<p>[/quote]</p>\n\n<p>..., unfortunately Lucas isn't telling us what KIND of single model it is. </p>\n\n<p>:-)</p>\n\n<p>[/quote]</p>\n\n<p>Haha! I think, going by the # of submissions of D&amp;L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)</p>\n\n<p>[/quote]</p>\n\n<p>Owen, I'm really interested on that too! :)  Just a bit of handcap.. Come on... \nAbout our number of submissions, its just that i like to do random tests to see if i'm not overfitting my cv somehow. My validation set works and we wouldnt need that many submissions to validate it. Just a hanful would be fine. Well not as few as owen, but not as much as we do.</p>",
      "rawMarkdown": "[quote=binga;86040]\r\n\r\n[quote=Owen;86014]\r\n\r\n[quote=Leustagos;85999]\r\n\r\nIts possible to get below 0.04100 with a single model...\r\n\r\n[/quote]\r\n\r\n..., unfortunately Lucas isn't telling us what KIND of single model it is. \r\n\r\n\r\n:-)\r\n\r\n[/quote]\r\n\r\nHaha! I think, going by the # of submissions of D&L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)\r\n\r\n[/quote]\r\n\r\nOwen, I'm really interested on that too! :)  Just a bit of handcap.. Come on... \r\nAbout our number of submissions, its just that i like to do random tests to see if i'm not overfitting my cv somehow. My validation set works and we wouldnt need that many submissions to validate it. Just a hanful would be fine. Well not as few as owen, but not as much as we do.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 85979,
      "author_name": "clustifier",
      "author_url": "",
      "post_date": "07/18/2015 20:44:40",
      "content": "<p>my current LB score (0.04546) is a single ftrl model with basic features from the most of the 8 data tables provided.</p>\n\n<p>What is yours?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85981,
      "author_name": "phanisrikanth",
      "author_url": "",
      "post_date": "07/18/2015 21:00:52",
      "content": "<p>Same here. 0.04562 - single FTRL. With raw features fro most of the tables.</p>\n\n<p>I am trying feature selection and interactions but the iteration time is really high. Takes almost 55 minutes for 1 epoch on 190M records.</p>\n\n<p>Any suggestions on how to approach interactions?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85985,
      "author_name": "triskelion",
      "author_url": "",
      "post_date": "07/18/2015 22:57:34",
      "content": "<blockquote>\n  <p>Any suggestions on how to approach interactions?</p>\n</blockquote>\n\n<p>Not competing yet, but try <a href=\"http://blogs.technet.com/b/machinelearning/archive/2014/09/24/online-learning-and-sub-linear-debugging.aspx\">sub-linear debugging</a>. </p>\n\n<p>Basically:</p>\n\n<ul>\n<li>try a run with a few interactions and see if the loss improves over 500k records. \n\n<blockquote>\n  <p>If so: cancel run, keep interactions and repeat, else: cancel run, discard interactions and repeat. </p>\n</blockquote></li>\n</ul>\n\n<p>Once you are done, make another run with all the interactions you found and confirm that it actually lowers progressive validation loss on complete train set. I think it may be considered a form of greedy forward feature interaction selection. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 85999,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/19/2015 01:49:33",
      "content": "<p>Its possible to get below 0.04100 with a single model...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86003,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "07/19/2015 02:23:23",
      "content": "<p>Below 0.043 with a single FTRL.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86004,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "07/19/2015 02:27:39",
      "content": "<p>I must miss something. LB tells me ads masters found some pretty straightforward way to get under 0.44. I hope I could figure it out before the deadline. This is the most fun part of kaggle after all :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86007,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "07/19/2015 02:34:32",
      "content": "<p>Why I can only get 0.050 with the raw features from 6 tables(no SisitsStream and PhoneRequestsStream)\nand the simple FTRL from Abhishek's code(without any changes).\nI am wondering if I need to do some parameter tunning or maybe my features are messed up.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86009,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "07/19/2015 03:01:28",
      "content": "<p>@SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86014,
      "author_name": "owenzhang1",
      "author_url": "",
      "post_date": "07/19/2015 03:18:01",
      "content": "<p>[quote=Leustagos;85999]</p>\n\n<p>Its possible to get below 0.04100 with a single model...</p>\n\n<p>[/quote]</p>\n\n<p>..., unfortunately Lucas isn't telling us what KIND of single model it is. </p>\n\n<p>:-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86025,
      "author_name": "remap1",
      "author_url": "",
      "post_date": "07/19/2015 05:17:06",
      "content": "<p>I'm still struggling getting a working validation set, which most would interpret as the set that produces similar scores to the LB. But after thinking about this, I think it's wrong to measure the quality of the validation set against the LB score. The reason is that if the scores are equal, it only guarantees that the validation set has a similar distribution of 1's versus 0's?</p>\n\n<p>With the LB score based on 30% of the data, if the current 30% is <em>not</em> representative of the entire LB set, then the LB position doesn't necessarily represent the winner. Also, it probably means that the best model/parameter set isn't necessarily a winning one.</p>\n\n<p>Are these observations correct and can this competition for that reason gives us some very interesting surprises at the end?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86026,
      "author_name": "skylibrary",
      "author_url": "",
      "post_date": "07/19/2015 05:30:05",
      "content": "<p>[quote=rcarson;86009]</p>\n\n<p>@SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P</p>\n\n<p>[/quote]\nThanks, rcarson:)\nI really should be more careful...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86040,
      "author_name": "phanisrikanth",
      "author_url": "",
      "post_date": "07/19/2015 08:47:03",
      "content": "<p>[quote=Owen;86014]</p>\n\n<p>[quote=Leustagos;85999]</p>\n\n<p>Its possible to get below 0.04100 with a single model...</p>\n\n<p>[/quote]</p>\n\n<p>..., unfortunately Lucas isn't telling us what KIND of single model it is. </p>\n\n<p>:-)</p>\n\n<p>[/quote]</p>\n\n<p>Haha! I think, going by the # of submissions of D&amp;L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 86069,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "07/19/2015 14:05:52",
      "content": "<p>[quote=binga;86040]</p>\n\n<p>[quote=Owen;86014]</p>\n\n<p>[quote=Leustagos;85999]</p>\n\n<p>Its possible to get below 0.04100 with a single model...</p>\n\n<p>[/quote]</p>\n\n<p>..., unfortunately Lucas isn't telling us what KIND of single model it is. </p>\n\n<p>:-)</p>\n\n<p>[/quote]</p>\n\n<p>Haha! I think, going by the # of submissions of D&amp;L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)</p>\n\n<p>[/quote]</p>\n\n<p>Owen, I'm really interested on that too! :)  Just a bit of handcap.. Come on... \nAbout our number of submissions, its just that i like to do random tests to see if i'm not overfitting my cv somehow. My validation set works and we wouldnt need that many submissions to validate it. Just a hanful would be fine. Well not as few as owen, but not as much as we do.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "85977": "What's your best single model score? We could push ourselves a little more if the leaders throw some light on their score :-)",
    "85979": "my current LB score (0.04546) is a single ftrl model with basic features from the most of the 8 data tables provided.\r\n\r\nWhat is yours?",
    "85981": "Same here. 0.04562 - single FTRL. With raw features fro most of the tables.\r\n\r\nI am trying feature selection and interactions but the iteration time is really high. Takes almost 55 minutes for 1 epoch on 190M records.\r\n\r\nAny suggestions on how to approach interactions?",
    "85985": ">Any suggestions on how to approach interactions?\r\n\r\nNot competing yet, but try [sub-linear debugging][1]. \r\n\r\nBasically:\r\n\r\n*  try a run with a few interactions and see if the loss improves over 500k records. \r\n>If so: cancel run, keep interactions and repeat, else: cancel run, discard interactions and repeat. \r\n\r\nOnce you are done, make another run with all the interactions you found and confirm that it actually lowers progressive validation loss on complete train set. I think it may be considered a form of greedy forward feature interaction selection. \r\n\r\n\r\n  [1]: http://blogs.technet.com/b/machinelearning/archive/2014/09/24/online-learning-and-sub-linear-debugging.aspx",
    "85999": "Its possible to get below 0.04100 with a single model...",
    "86003": "Below 0.043 with a single FTRL.",
    "86004": "I must miss something. LB tells me ads masters found some pretty straightforward way to get under 0.44. I hope I could figure it out before the deadline. This is the most fun part of kaggle after all :D",
    "86007": "Why I can only get 0.050 with the raw features from 6 tables(no SisitsStream and PhoneRequestsStream)\r\nand the simple FTRL from Abhishek's code(without any changes).\r\nI am wondering if I need to do some parameter tunning or maybe my features are messed up.",
    "86009": "SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P",
    "86014": "[quote=Leustagos;85999]\r\n\r\nIts possible to get below 0.04100 with a single model...\r\n\r\n[/quote]\r\n\r\n..., unfortunately Lucas isn't telling us what KIND of single model it is. \r\n\r\n\r\n:-)",
    "86025": "I'm still struggling getting a working validation set, which most would interpret as the set that produces similar scores to the LB. But after thinking about this, I think it's wrong to measure the quality of the validation set against the LB score. The reason is that if the scores are equal, it only guarantees that the validation set has a similar distribution of 1's versus 0's?\r\n\r\nWith the LB score based on 30% of the data, if the current 30% is *not* representative of the entire LB set, then the LB position doesn't necessarily represent the winner. Also, it probably means that the best model/parameter set isn't necessarily a winning one.\r\n\r\nAre these observations correct and can this competition for that reason gives us some very interesting surprises at the end?",
    "86026": "[quote=rcarson;86009]\r\n\r\n@SkyLibrary, I assumed you didn't change bits from 20 to 2**20 :P\r\n\r\n[/quote]\r\nThanks, rcarson:)\r\nI really should be more careful...",
    "86040": "[quote=Owen;86014]\r\n\r\n[quote=Leustagos;85999]\r\n\r\nIts possible to get below 0.04100 with a single model...\r\n\r\n[/quote]\r\n\r\n..., unfortunately Lucas isn't telling us what KIND of single model it is. \r\n\r\n\r\n:-)\r\n\r\n[/quote]\r\n\r\nHaha! I think, going by the # of submissions of D&L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)",
    "86069": "[quote=binga;86040]\r\n\r\n[quote=Owen;86014]\r\n\r\n[quote=Leustagos;85999]\r\n\r\nIts possible to get below 0.04100 with a single model...\r\n\r\n[/quote]\r\n\r\n..., unfortunately Lucas isn't telling us what KIND of single model it is. \r\n\r\n\r\n:-)\r\n\r\n[/quote]\r\n\r\nHaha! I think, going by the # of submissions of D&L, they must have done some fair amount of feature engineering which is really helping them. (My hunch) I guess, you should tell us more what MODEL you are using, since you are sitting pretty at the top with just 2 frikkin submissions! ;-)\r\n\r\n[/quote]\r\n\r\nOwen, I'm really interested on that too! :)  Just a bit of handcap.. Come on... \r\nAbout our number of submissions, its just that i like to do random tests to see if i'm not overfitting my cv somehow. My validation set works and we wouldnt need that many submissions to validate it. Just a hanful would be fine. Well not as few as owen, but not as much as we do."
  },
  "source": "meta"
}