{
  "id": 24111,
  "title": "is it just me?",
  "url": "/competitions/outbrain-click-prediction/discussion/24111",
  "author_name": "raddar",
  "post_date": "2016-10-05T20:40:07.267000",
  "votes": 37,
  "comment_count": 19,
  "views": 3297,
  "content": "<p>Am I the only one noticing strong negative correlation with prize pool and data size in recent competitions ?:)</p>\n\n<p>Anyway, this competition looks interesting, but I don't think will invest my RAM into this. Good luck everyone!</p>",
  "messages": [
    {
      "id": 137879,
      "postDate": "2016-10-05T20:40:07.267Z",
      "content": "<p>Am I the only one noticing strong negative correlation with prize pool and data size in recent competitions ?:)</p>\n\n<p>Anyway, this competition looks interesting, but I don't think will invest my RAM into this. Good luck everyone!</p>",
      "rawMarkdown": "Am I the only one noticing strong negative correlation with prize pool and data size in recent competitions ?:)\r\n\r\nAnyway, this competition looks interesting, but I don't think will invest my RAM into this. Good luck everyone!",
      "votes": 37
    },
    {
      "id": 138091,
      "postDate": "2016-10-07T05:49:23.830Z",
      "content": "<p>[quote=Giulio;138042]</p>\n\n<p>I might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is &quot;reverting to the mean&quot;, in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).</p>\n\n<p>[/quote]</p>\n\n<p>I don't think it is that the data is scrutinized, but I think companies are realizing that implementing the solutions is challenging. While it seems like you get free labor, the company may not be able to understand how to use it. Kaggle solutions are complex, and probably more accurate than the business needs. Also a lot of the work is redundant, and it takes time to prepare and run competitions. The prize money is usually a small amount anyway compared with the time, skill and luck it takes to win a competition.</p>\n\n<p>On another note, I really wish Kaggle would strongly discourage sponsors from posting datasets of this size. There is really no reason why the admins could not of scaled this dataset down - percent statistical uncertainty goes like 1 / sqrt(N), so this dataset could be much smaller and still be dominated by systematic uncertainty. Of course I can scale the dataset down myself, but when it won't even fit on my machine this is nontrivial. While big data presents interesting technical challenges, it requires dedicated infrastructure to tackle. Finally, in terms of public relations, companies would be better served with a smaller dataset that allows more people to try the competition out.</p>",
      "rawMarkdown": "[quote=Giulio;138042]\r\n\r\nI might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is \"reverting to the mean\", in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).\r\n\r\n[/quote]\r\n\r\nI don't think it is that the data is scrutinized, but I think companies are realizing that implementing the solutions is challenging. While it seems like you get free labor, the company may not be able to understand how to use it. Kaggle solutions are complex, and probably more accurate than the business needs. Also a lot of the work is redundant, and it takes time to prepare and run competitions. The prize money is usually a small amount anyway compared with the time, skill and luck it takes to win a competition.\r\n\r\nOn another note, I really wish Kaggle would strongly discourage sponsors from posting datasets of this size. There is really no reason why the admins could not of scaled this dataset down - percent statistical uncertainty goes like 1 / sqrt(N), so this dataset could be much smaller and still be dominated by systematic uncertainty. Of course I can scale the dataset down myself, but when it won't even fit on my machine this is nontrivial. While big data presents interesting technical challenges, it requires dedicated infrastructure to tackle. Finally, in terms of public relations, companies would be better served with a smaller dataset that allows more people to try the competition out.",
      "votes": 20
    },
    {
      "id": 137910,
      "postDate": "2016-10-05T23:25:16.463Z",
      "content": "<p>[quote=raddar;137879]</p>\n\n<p>but I don't think will invest my RAM into this. Good luck everyone!</p>\n\n<p>[/quote]</p>\n\n<p>Thank you ! That's good news to me! I am always luckier with big dataset when top kagglers are not interested.</p>",
      "rawMarkdown": "[quote=raddar;137879]\r\n\r\nbut I don't think will invest my RAM into this. Good luck everyone!\r\n\r\n[/quote]\r\n\r\nThank you ! That's good news to me! I am always luckier with big dataset when top kagglers are not interested.",
      "votes": 15
    },
    {
      "id": 137908,
      "postDate": "2016-10-05T23:04:09.977Z",
      "content": "<p>I am starting to have a sinking feeling that the kinds of skills that the Kaggle competition problems are designed for are rapidly becoming commoditized. It seems that the skills required to get the best predictive model for a laptop-sized datasets can easily be found inhouse at most organizations. Only the really large and/or messy datasets are &quot;outsourced&quot; to Kaggle, there are fewer and fewer of those, and the prize money is going down. :/</p>",
      "rawMarkdown": "I am starting to have a sinking feeling that the kinds of skills that the Kaggle competition problems are designed for are rapidly becoming commoditized. It seems that the skills required to get the best predictive model for a laptop-sized datasets can easily be found inhouse at most organizations. Only the really large and/or messy datasets are \"outsourced\" to Kaggle, there are fewer and fewer of those, and the prize money is going down. :/",
      "votes": 13
    },
    {
      "id": 137986,
      "postDate": "2016-10-06T13:30:23.557Z",
      "content": "<p>What?!</p>\n\n<p>I'm came here to learn about celeb tricks for losing belly fat.</p>",
      "rawMarkdown": "What?!\r\n\r\nI'm came here to learn about celeb tricks for losing belly fat.",
      "votes": 14
    },
    {
      "id": 138042,
      "postDate": "2016-10-06T19:04:31.047Z",
      "content": "<p>I might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is &quot;reverting to the mean&quot;, in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).</p>\n\n<p>Again, no data on this, but just from talking to people and sponsors the general feeling is that many Kaggle competitions over the past couple of years had mostly a PR connotation. </p>\n\n<p>If that is true, anyone with any degree of judgement starts doing the math. What is a PR stunt on Kaggle worth? Let's say is X. Now, given that many competition end up reveling some degree of problem (leakage being the worst) that essentially backfires and turns a PR stunt in a PR nightmare, what is a Kaggle competition now worth? The window for a positive (or acceptable) ROI becomes smaller and smaller and the only way to mitigate that risk is either with lower investments (aka prizes) or by not sponsoring a competition. </p>\n\n<p>EDIT: all of that being said, Kaggle is very seasonal as it is essentially correlated to budgets and their seasonality. So, I wouldn't be surprised to see more competitions and larger prizes as we enter a new year.</p>",
      "rawMarkdown": "I might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is \"reverting to the mean\", in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).\r\n\r\nAgain, no data on this, but just from talking to people and sponsors the general feeling is that many Kaggle competitions over the past couple of years had mostly a PR connotation. \r\n\r\nIf that is true, anyone with any degree of judgement starts doing the math. What is a PR stunt on Kaggle worth? Let's say is X. Now, given that many competition end up reveling some degree of problem (leakage being the worst) that essentially backfires and turns a PR stunt in a PR nightmare, what is a Kaggle competition now worth? The window for a positive (or acceptable) ROI becomes smaller and smaller and the only way to mitigate that risk is either with lower investments (aka prizes) or by not sponsoring a competition. \r\n\r\nEDIT: all of that being said, Kaggle is very seasonal as it is essentially correlated to budgets and their seasonality. So, I wouldn't be surprised to see more competitions and larger prizes as we enter a new year.",
      "votes": 10
    },
    {
      "id": 137898,
      "postDate": "2016-10-05T21:44:31.673Z",
      "content": "<p>Yes, that is how (that/this?)business work, they will exploit you as much as you let them ;)</p>",
      "rawMarkdown": "Yes, that is how (that/this?)business work, they will exploit you as much as you let them ;)",
      "votes": 9
    },
    {
      "id": 138204,
      "postDate": "2016-10-07T16:34:32.847Z",
      "content": "<p>[quote=Fernando Constantino;138198]</p>\n\n<p>Bright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. </p>\n\n<p>[/quote]</p>\n\n<p>LOL, I'm just guessing here, but I would venture that you haven't been around online discussion boards all that much. ;) This is as polite, respectful, and most importantly CONSTRUCTIVE criticism as anything I have ever come across. </p>",
      "rawMarkdown": "[quote=Fernando Constantino;138198]\r\n\r\nBright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. \r\n\r\n[/quote]\r\n\r\nLOL, I'm just guessing here, but I would venture that you haven't been around online discussion boards all that much. ;) This is as polite, respectful, and most importantly CONSTRUCTIVE criticism as anything I have ever come across. ",
      "votes": 8
    },
    {
      "id": 137975,
      "postDate": "2016-10-06T11:16:14.103Z",
      "content": "<p>I believe this can be attributed to a lot of things, one of them being the fact that most of the companies don't use the models they obtain from kaggle winning solutions. </p>\n\n<p>When they aren't gonna use the models anyway, why pay a hefty prize money for it. The competition amongst the top kagglers and the creativity all participants bring to the table is enough for getting thoughtful insight about the data, something a team of data scientists working in-house under time/target constraints might not be able to come across. Not because of limited skillset, but simply because of limited time. </p>\n\n<p>50000$ bucks for a palatte of around 2000 people (talking about the Redhat competiton) (I doubt competition ever goes to that level, considering the constraints in question) of different backgrounds, both educational and geographical, literate enough to work their way around data, with the top 10-15% really exceptional at their craft, working for a duration of 2-3 months! Even if we only consider the top 200 candidates to be the ones that give substantial contributions, the cost turns out to be 250$ per &quot;data scientist&quot; for 2-3 months (whatever the duration might be). This is a goddarn bargain! Any company would be smacking their lips!</p>\n\n<p>I believe this trend is here to stay.</p>",
      "rawMarkdown": "I believe this can be attributed to a lot of things, one of them being the fact that most of the companies don't use the models they obtain from kaggle winning solutions. \r\n\r\nWhen they aren't gonna use the models anyway, why pay a hefty prize money for it. The competition amongst the top kagglers and the creativity all participants bring to the table is enough for getting thoughtful insight about the data, something a team of data scientists working in-house under time/target constraints might not be able to come across. Not because of limited skillset, but simply because of limited time. \r\n\r\n50000$ bucks for a palatte of around 2000 people (talking about the Redhat competiton) (I doubt competition ever goes to that level, considering the constraints in question) of different backgrounds, both educational and geographical, literate enough to work their way around data, with the top 10-15% really exceptional at their craft, working for a duration of 2-3 months! Even if we only consider the top 200 candidates to be the ones that give substantial contributions, the cost turns out to be 250$ per \"data scientist\" for 2-3 months (whatever the duration might be). This is a goddarn bargain! Any company would be smacking their lips!\r\n\r\nI believe this trend is here to stay.",
      "votes": 6
    },
    {
      "id": 138186,
      "postDate": "2016-10-07T15:34:56.333Z",
      "content": "<p>With regards to dataset size, I was hoping this one would be my first Kaggle submission, but it has seriously put me off.</p>\n\n<p>I work a full time DS job and was looking for a more recreational predictive modelling exercise during the limited play-time I have available - and also learn a thing or too from this great community.</p>\n\n<p>But I deal with large, problematic datasets on a daily basis - I'd rather spend time studying recommendation engines than optimally subsetting this monster.. my machine barely has the disk space for unzipping it!</p>\n\n<p>A bit disappointed here.</p>",
      "rawMarkdown": "With regards to dataset size, I was hoping this one would be my first Kaggle submission, but it has seriously put me off.\r\n\r\nI work a full time DS job and was looking for a more recreational predictive modelling exercise during the limited play-time I have available - and also learn a thing or too from this great community.\r\n\r\nBut I deal with large, problematic datasets on a daily basis - I'd rather spend time studying recommendation engines than optimally subsetting this monster.. my machine barely has the disk space for unzipping it!\r\n\r\nA bit disappointed here.",
      "votes": 4
    },
    {
      "id": 138193,
      "postDate": "2016-10-07T16:03:03.450Z",
      "content": "<p>[quote=Fernando Constantino;138191]</p>\n\n<p>I find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..</p>\n\n<p>[/quote]</p>\n\n<p>Are u serious? So you want bright and intelligent people not to say whats on their mind...</p>",
      "rawMarkdown": "[quote=Fernando Constantino;138191]\r\n\r\nI find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\r\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..\r\n\r\n\r\n\r\n[/quote]\r\n\r\nAre u serious? So you want bright and intelligent people not to say whats on their mind...",
      "votes": 3
    },
    {
      "id": 137957,
      "postDate": "2016-10-06T09:06:25.863Z",
      "content": "<p>Perhaps Kaggle is being used as a benchmarker?</p>",
      "rawMarkdown": "Perhaps Kaggle is being used as a benchmarker?",
      "votes": 3
    },
    {
      "id": 137953,
      "postDate": "2016-10-06T08:52:58.367Z",
      "content": "<p>@Bojan\nI don't think many companies actually deploy the winning algorithms in production but more see it as a way to data mine ideas and knowledge. That prizes go down is only demand/supply like you said and if a company has a lot of data they surely have a couple data scientists already in house.</p>",
      "rawMarkdown": "@Bojan\r\nI don't think many companies actually deploy the winning algorithms in production but more see it as a way to data mine ideas and knowledge. That prizes go down is only demand/supply like you said and if a company has a lot of data they surely have a couple data scientists already in house.",
      "votes": 4
    },
    {
      "id": 137990,
      "postDate": "2016-10-06T13:57:44.240Z",
      "content": "<p>@Scirpus they are called deep fat destroyers these days.</p>",
      "rawMarkdown": "@Scirpus they are called deep fat destroyers these days.",
      "votes": 2
    },
    {
      "id": 137989,
      "postDate": "2016-10-06T13:55:14.827Z",
      "content": "<p>@Scripus I almost missed this, hilarious!</p>",
      "rawMarkdown": "@Scripus I almost missed this, hilarious!",
      "votes": 2
    },
    {
      "id": 138217,
      "postDate": "2016-10-07T16:59:48.767Z",
      "content": "<p>Just to clarify. Prior to 'LOLing' at my post, all your comments were polite and constructive, as was Raddar's and some others. I was only referring to some of the comments in this thread. But given my huge success at keeping things in a good vibe, I will stick to coding for a while ;)</p>",
      "rawMarkdown": "Just to clarify. Prior to 'LOLing' at my post, all your comments were polite and constructive, as was Raddar's and some others. I was only referring to some of the comments in this thread. But given my huge success at keeping things in a good vibe, I will stick to coding for a while ;)\r\n"
    },
    {
      "id": 138198,
      "postDate": "2016-10-07T16:10:06.720Z",
      "content": "<p>Bright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. Anyway, I just wanted to lower the tone of this thread.</p>",
      "rawMarkdown": "Bright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. Anyway, I just wanted to lower the tone of this thread.",
      "votes": -6
    },
    {
      "id": 138191,
      "postDate": "2016-10-07T15:51:36Z",
      "content": "<p>I find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..</p>",
      "rawMarkdown": "I find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\r\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..\r\n\r\n",
      "votes": -6
    },
    {
      "id": 138283,
      "postDate": "2016-10-07T22:29:33.470Z",
      "content": "<p>If you look at the competitions that are held one or two years ago, the average prize is even lower...</p>",
      "rawMarkdown": "If you look at the competitions that are held one or two years ago, the average prize is even lower..."
    },
    {
      "id": 137882,
      "postDate": "2016-10-05T20:45:06.160Z",
      "content": "<p>Well it's Outbrain, which has a strong negative correlation between how they present themselves and how horrible they actually are :)</p>",
      "rawMarkdown": "Well it's Outbrain, which has a strong negative correlation between how they present themselves and how horrible they actually are :)",
      "votes": 13,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 138091,
      "author_name": "Jared Turkewitz",
      "author_url": "",
      "post_date": "2016-10-07T05:49:23.830000",
      "content": "<p>[quote=Giulio;138042]</p>\n\n<p>I might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is &quot;reverting to the mean&quot;, in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).</p>\n\n<p>[/quote]</p>\n\n<p>I don't think it is that the data is scrutinized, but I think companies are realizing that implementing the solutions is challenging. While it seems like you get free labor, the company may not be able to understand how to use it. Kaggle solutions are complex, and probably more accurate than the business needs. Also a lot of the work is redundant, and it takes time to prepare and run competitions. The prize money is usually a small amount anyway compared with the time, skill and luck it takes to win a competition.</p>\n\n<p>On another note, I really wish Kaggle would strongly discourage sponsors from posting datasets of this size. There is really no reason why the admins could not of scaled this dataset down - percent statistical uncertainty goes like 1 / sqrt(N), so this dataset could be much smaller and still be dominated by systematic uncertainty. Of course I can scale the dataset down myself, but when it won't even fit on my machine this is nontrivial. While big data presents interesting technical challenges, it requires dedicated infrastructure to tackle. Finally, in terms of public relations, companies would be better served with a smaller dataset that allows more people to try the competition out.</p>",
      "votes": 20,
      "replies": []
    },
    {
      "id": 137910,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2016-10-05T23:25:16.463000",
      "content": "<p>[quote=raddar;137879]</p>\n\n<p>but I don't think will invest my RAM into this. Good luck everyone!</p>\n\n<p>[/quote]</p>\n\n<p>Thank you ! That's good news to me! I am always luckier with big dataset when top kagglers are not interested.</p>",
      "votes": 15,
      "replies": []
    },
    {
      "id": 137908,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2016-10-05T23:04:09.977000",
      "content": "<p>I am starting to have a sinking feeling that the kinds of skills that the Kaggle competition problems are designed for are rapidly becoming commoditized. It seems that the skills required to get the best predictive model for a laptop-sized datasets can easily be found inhouse at most organizations. Only the really large and/or messy datasets are &quot;outsourced&quot; to Kaggle, there are fewer and fewer of those, and the prize money is going down. :/</p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 137986,
      "author_name": "Scirpus",
      "author_url": "",
      "post_date": "2016-10-06T13:30:23.557000",
      "content": "<p>What?!</p>\n\n<p>I'm came here to learn about celeb tricks for losing belly fat.</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 138042,
      "author_name": "Giulio",
      "author_url": "",
      "post_date": "2016-10-06T19:04:31.047000",
      "content": "<p>I might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is &quot;reverting to the mean&quot;, in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).</p>\n\n<p>Again, no data on this, but just from talking to people and sponsors the general feeling is that many Kaggle competitions over the past couple of years had mostly a PR connotation. </p>\n\n<p>If that is true, anyone with any degree of judgement starts doing the math. What is a PR stunt on Kaggle worth? Let's say is X. Now, given that many competition end up reveling some degree of problem (leakage being the worst) that essentially backfires and turns a PR stunt in a PR nightmare, what is a Kaggle competition now worth? The window for a positive (or acceptable) ROI becomes smaller and smaller and the only way to mitigate that risk is either with lower investments (aka prizes) or by not sponsoring a competition. </p>\n\n<p>EDIT: all of that being said, Kaggle is very seasonal as it is essentially correlated to budgets and their seasonality. So, I wouldn't be surprised to see more competitions and larger prizes as we enter a new year.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 137898,
      "author_name": "NxGTR",
      "author_url": "",
      "post_date": "2016-10-05T21:44:31.673000",
      "content": "<p>Yes, that is how (that/this?)business work, they will exploit you as much as you let them ;)</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 138204,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2016-10-07T16:34:32.847000",
      "content": "<p>[quote=Fernando Constantino;138198]</p>\n\n<p>Bright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. </p>\n\n<p>[/quote]</p>\n\n<p>LOL, I'm just guessing here, but I would venture that you haven't been around online discussion boards all that much. ;) This is as polite, respectful, and most importantly CONSTRUCTIVE criticism as anything I have ever come across. </p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 137975,
      "author_name": "SarthakYadav",
      "author_url": "",
      "post_date": "2016-10-06T11:16:14.103000",
      "content": "<p>I believe this can be attributed to a lot of things, one of them being the fact that most of the companies don't use the models they obtain from kaggle winning solutions. </p>\n\n<p>When they aren't gonna use the models anyway, why pay a hefty prize money for it. The competition amongst the top kagglers and the creativity all participants bring to the table is enough for getting thoughtful insight about the data, something a team of data scientists working in-house under time/target constraints might not be able to come across. Not because of limited skillset, but simply because of limited time. </p>\n\n<p>50000$ bucks for a palatte of around 2000 people (talking about the Redhat competiton) (I doubt competition ever goes to that level, considering the constraints in question) of different backgrounds, both educational and geographical, literate enough to work their way around data, with the top 10-15% really exceptional at their craft, working for a duration of 2-3 months! Even if we only consider the top 200 candidates to be the ones that give substantial contributions, the cost turns out to be 250$ per &quot;data scientist&quot; for 2-3 months (whatever the duration might be). This is a goddarn bargain! Any company would be smacking their lips!</p>\n\n<p>I believe this trend is here to stay.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 138186,
      "author_name": "Alex Spanos",
      "author_url": "",
      "post_date": "2016-10-07T15:34:56.333000",
      "content": "<p>With regards to dataset size, I was hoping this one would be my first Kaggle submission, but it has seriously put me off.</p>\n\n<p>I work a full time DS job and was looking for a more recreational predictive modelling exercise during the limited play-time I have available - and also learn a thing or too from this great community.</p>\n\n<p>But I deal with large, problematic datasets on a daily basis - I'd rather spend time studying recommendation engines than optimally subsetting this monster.. my machine barely has the disk space for unzipping it!</p>\n\n<p>A bit disappointed here.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 138193,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-10-07T16:03:03.450000",
      "content": "<p>[quote=Fernando Constantino;138191]</p>\n\n<p>I find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..</p>\n\n<p>[/quote]</p>\n\n<p>Are u serious? So you want bright and intelligent people not to say whats on their mind...</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 137957,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2016-10-06T09:06:25.863000",
      "content": "<p>Perhaps Kaggle is being used as a benchmarker?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 137953,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-10-06T08:52:58.367000",
      "content": "<p>@Bojan\nI don't think many companies actually deploy the winning algorithms in production but more see it as a way to data mine ideas and knowledge. That prizes go down is only demand/supply like you said and if a company has a lot of data they surely have a couple data scientists already in house.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 137990,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-10-06T13:57:44.240000",
      "content": "<p>@Scirpus they are called deep fat destroyers these days.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 137989,
      "author_name": "timmy",
      "author_url": "",
      "post_date": "2016-10-06T13:55:14.827000",
      "content": "<p>@Scripus I almost missed this, hilarious!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 138217,
      "author_name": "Fernando Constantino",
      "author_url": "",
      "post_date": "2016-10-07T16:59:48.767000",
      "content": "<p>Just to clarify. Prior to 'LOLing' at my post, all your comments were polite and constructive, as was Raddar's and some others. I was only referring to some of the comments in this thread. But given my huge success at keeping things in a good vibe, I will stick to coding for a while ;)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 138198,
      "author_name": "Fernando Constantino",
      "author_url": "",
      "post_date": "2016-10-07T16:10:06.720000",
      "content": "<p>Bright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. Anyway, I just wanted to lower the tone of this thread.</p>",
      "votes": -6,
      "replies": []
    },
    {
      "id": 138191,
      "author_name": "Fernando Constantino",
      "author_url": "",
      "post_date": "2016-10-07T15:51:36",
      "content": "<p>I find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..</p>",
      "votes": -6,
      "replies": []
    },
    {
      "id": 138283,
      "author_name": "Jiahuan He",
      "author_url": "",
      "post_date": "2016-10-07T22:29:33.470000",
      "content": "<p>If you look at the competitions that are held one or two years ago, the average prize is even lower...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 137882,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-10-05T20:45:06.160000",
      "content": "<p>Well it's Outbrain, which has a strong negative correlation between how they present themselves and how horrible they actually are :)</p>",
      "votes": 13,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "137879": "Am I the only one noticing strong negative correlation with prize pool and data size in recent competitions ?:)\r\n\r\nAnyway, this competition looks interesting, but I don't think will invest my RAM into this. Good luck everyone!",
    "138091": "[quote=Giulio;138042]\r\n\r\nI might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is \"reverting to the mean\", in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).\r\n\r\n[/quote]\r\n\r\nI don't think it is that the data is scrutinized, but I think companies are realizing that implementing the solutions is challenging. While it seems like you get free labor, the company may not be able to understand how to use it. Kaggle solutions are complex, and probably more accurate than the business needs. Also a lot of the work is redundant, and it takes time to prepare and run competitions. The prize money is usually a small amount anyway compared with the time, skill and luck it takes to win a competition.\r\n\r\nOn another note, I really wish Kaggle would strongly discourage sponsors from posting datasets of this size. There is really no reason why the admins could not of scaled this dataset down - percent statistical uncertainty goes like 1 / sqrt(N), so this dataset could be much smaller and still be dominated by systematic uncertainty. Of course I can scale the dataset down myself, but when it won't even fit on my machine this is nontrivial. While big data presents interesting technical challenges, it requires dedicated infrastructure to tackle. Finally, in terms of public relations, companies would be better served with a smaller dataset that allows more people to try the competition out.",
    "137910": "[quote=raddar;137879]\r\n\r\nbut I don't think will invest my RAM into this. Good luck everyone!\r\n\r\n[/quote]\r\n\r\nThank you ! That's good news to me! I am always luckier with big dataset when top kagglers are not interested.",
    "137908": "I am starting to have a sinking feeling that the kinds of skills that the Kaggle competition problems are designed for are rapidly becoming commoditized. It seems that the skills required to get the best predictive model for a laptop-sized datasets can easily be found inhouse at most organizations. Only the really large and/or messy datasets are \"outsourced\" to Kaggle, there are fewer and fewer of those, and the prize money is going down. :/",
    "137986": "What?!\r\n\r\nI'm came here to learn about celeb tricks for losing belly fat.",
    "138042": "I might very well be wrong on this but I have the  feeling that Kaggle has been exiting the Cinderella/novelty period over the last year and entering a different portion in its journey to maturity. In particular I think the sponsor market has really been settling down on the actual market value of running a Kaggle competition. The market value of croudsourcing is \"reverting to the mean\", in the sense that what 5 years ago was a no brainier (get hundreds of people to solve a problem for a modest financial investment), today becomes at least debatable because of the in-depth scrutiny your data goes through when exposed to hundreds of people who collectively will spend very significant amount of time just trying to find errors in the data to exploit (and BTW- these are now people with 5+ years of experience doing so).\r\n\r\nAgain, no data on this, but just from talking to people and sponsors the general feeling is that many Kaggle competitions over the past couple of years had mostly a PR connotation. \r\n\r\nIf that is true, anyone with any degree of judgement starts doing the math. What is a PR stunt on Kaggle worth? Let's say is X. Now, given that many competition end up reveling some degree of problem (leakage being the worst) that essentially backfires and turns a PR stunt in a PR nightmare, what is a Kaggle competition now worth? The window for a positive (or acceptable) ROI becomes smaller and smaller and the only way to mitigate that risk is either with lower investments (aka prizes) or by not sponsoring a competition. \r\n\r\nEDIT: all of that being said, Kaggle is very seasonal as it is essentially correlated to budgets and their seasonality. So, I wouldn't be surprised to see more competitions and larger prizes as we enter a new year.",
    "137898": "Yes, that is how (that/this?)business work, they will exploit you as much as you let them ;)",
    "138204": "[quote=Fernando Constantino;138198]\r\n\r\nBright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. \r\n\r\n[/quote]\r\n\r\nLOL, I'm just guessing here, but I would venture that you haven't been around online discussion boards all that much. ;) This is as polite, respectful, and most importantly CONSTRUCTIVE criticism as anything I have ever come across. ",
    "137975": "I believe this can be attributed to a lot of things, one of them being the fact that most of the companies don't use the models they obtain from kaggle winning solutions. \r\n\r\nWhen they aren't gonna use the models anyway, why pay a hefty prize money for it. The competition amongst the top kagglers and the creativity all participants bring to the table is enough for getting thoughtful insight about the data, something a team of data scientists working in-house under time/target constraints might not be able to come across. Not because of limited skillset, but simply because of limited time. \r\n\r\n50000$ bucks for a palatte of around 2000 people (talking about the Redhat competiton) (I doubt competition ever goes to that level, considering the constraints in question) of different backgrounds, both educational and geographical, literate enough to work their way around data, with the top 10-15% really exceptional at their craft, working for a duration of 2-3 months! Even if we only consider the top 200 candidates to be the ones that give substantial contributions, the cost turns out to be 250$ per \"data scientist\" for 2-3 months (whatever the duration might be). This is a goddarn bargain! Any company would be smacking their lips!\r\n\r\nI believe this trend is here to stay.",
    "138186": "With regards to dataset size, I was hoping this one would be my first Kaggle submission, but it has seriously put me off.\r\n\r\nI work a full time DS job and was looking for a more recreational predictive modelling exercise during the limited play-time I have available - and also learn a thing or too from this great community.\r\n\r\nBut I deal with large, problematic datasets on a daily basis - I'd rather spend time studying recommendation engines than optimally subsetting this monster.. my machine barely has the disk space for unzipping it!\r\n\r\nA bit disappointed here.",
    "138193": "[quote=Fernando Constantino;138191]\r\n\r\nI find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\r\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..\r\n\r\n\r\n\r\n[/quote]\r\n\r\nAre u serious? So you want bright and intelligent people not to say whats on their mind...",
    "137957": "Perhaps Kaggle is being used as a benchmarker?",
    "137953": "@Bojan\r\nI don't think many companies actually deploy the winning algorithms in production but more see it as a way to data mine ideas and knowledge. That prizes go down is only demand/supply like you said and if a company has a lot of data they surely have a couple data scientists already in house.",
    "137990": "@Scirpus they are called deep fat destroyers these days.",
    "137989": "@Scripus I almost missed this, hilarious!",
    "138217": "Just to clarify. Prior to 'LOLing' at my post, all your comments were polite and constructive, as was Raddar's and some others. I was only referring to some of the comments in this thread. But given my huge success at keeping things in a good vibe, I will stick to coding for a while ;)\r\n",
    "138198": "Bright and intelligent people need to be polite when expressing their opinions. Some of the posts in this topic are not. Anyway, I just wanted to lower the tone of this thread.",
    "138191": "I find this post to be very discouraging and borderline disrespectful with the sponsor. As anybody else, I would like to see a higher prize and less RAM and CPU intensive competitions, but this may not fit with the sponsor's business needs.\r\nSo, let's calm down, and most importantly, transform that downvoting impulse you may have now into something better such as uploading a nice performing kernel (which may save me some time to start this competition)..\r\n\r\n",
    "138283": "If you look at the competitions that are held one or two years ago, the average prize is even lower...",
    "137882": "Well it's Outbrain, which has a strong negative correlation between how they present themselves and how horrible they actually are :)"
  }
}