{
  "id": 22016,
  "title": "What is Kaggle's / Avito's opinion on precomputed image hashes?",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/22016",
  "author_name": "",
  "post_date": "2016-07-02T22:38:18.917Z",
  "votes": 1,
  "comment_count": 10,
  "views": 1494,
  "content": "<p>There are multiple precomputed image hash dumps on the forums. What will Kaggle's / Avito's stance be, if a prize eligible team were using these dumps as features?</p>",
  "messages": [
    {
      "id": "125809",
      "postDate": "07/02/2016 22:38:18",
      "content": "<p>There are multiple precomputed image hash dumps on the forums. What will Kaggle's / Avito's stance be, if a prize eligible team were using these dumps as features?</p>",
      "rawMarkdown": "There are multiple precomputed image hash dumps on the forums. What will Kaggle's / Avito's stance be, if a prize eligible team were using these dumps as features?",
      "votes": null
    },
    {
      "id": "125822",
      "postDate": "07/03/2016 01:29:58",
      "content": "<p>What rule do you think is being violated?</p>",
      "rawMarkdown": "What rule do you think is being violated?",
      "votes": null
    },
    {
      "id": "125827",
      "postDate": "07/03/2016 05:00:55",
      "content": "<p>inversion, the problem is that results won't be reproducable.</p>\n\n<p>But if somebody using those precomputed hashes wins the prize, they can always write a script to compute them.</p>",
      "rawMarkdown": "inversion, the problem is that results won't be reproducable.\r\n\r\nBut if somebody using those precomputed hashes wins the prize, they can always write a script to compute them.",
      "votes": null
    },
    {
      "id": "125835",
      "postDate": "07/03/2016 08:15:06",
      "content": "<p>[quote=inversion;125822]</p>\n\n<p>What rule do you think is being violated?</p>\n\n<p>[/quote]</p>\n\n<p>I was thinking of common sense, rather than a specific rule. But yeah, let's say reproducibility.</p>\n\n<p>How is this any different than, say, the first place team dumping the probability outputs of their best model? That doesn't violate any specific rules. Would Kaggle / Avito / everyone be okay with using those probs as features?</p>\n\n<p>That being said, it could just be me. I still see these prize competitions as competitive events for a prize, rather than let's-all-get-together-and-see-what-we-come-up-with.</p>",
      "rawMarkdown": "[quote=inversion;125822]\r\n\r\nWhat rule do you think is being violated?\r\n\r\n[/quote]\r\n\r\nI was thinking of common sense, rather than a specific rule. But yeah, let's say reproducibility.\r\n\r\nHow is this any different than, say, the first place team dumping the probability outputs of their best model? That doesn't violate any specific rules. Would Kaggle / Avito / everyone be okay with using those probs as features?\r\n\r\nThat being said, it could just be me. I still see these prize competitions as competitive events for a prize, rather than let's-all-get-together-and-see-what-we-come-up-with.",
      "votes": null
    },
    {
      "id": "125867",
      "postDate": "07/03/2016 20:08:47",
      "content": "<p>@barisumog - I do see where you are coming from, and initially grimaced at the request as well as the fulfillment  of the pre-computed image hashes. The way I figure it, though, is that these are relatively straight forward calculations, and thus don't necessarily provide any real competitive advantage, especially since scripts have already been posted showing how to do it. In my opinion, this is not much different than posting the image height and width.</p>",
      "rawMarkdown": "barisumog - I do see where you are coming from, and initially grimaced at the request as well as the fulfillment  of the pre-computed image hashes. The way I figure it, though, is that these are relatively straight forward calculations, and thus don't necessarily provide any real competitive advantage, especially since scripts have already been posted showing how to do it. In my opinion, this is not much different than posting the image height and width.",
      "votes": null
    },
    {
      "id": "125876",
      "postDate": "07/03/2016 23:41:07",
      "content": "<p>I agree it might not be a huge problem in this case. It would still be unfair, if someone spent money on bandwidth and hardware to download the data and calculate the features early on, and in the last couple weeks they are just dumped here.</p>\n\n<p>What is more concerning imho, is if this becomes commonplace. In the very least, it escalates the current problems with scripts to a higher level.</p>\n\n<p>And to add to that, I could use commercial software to calculate some decisive features and post them here. Or I could use open source tools, but manually tweak 5% of the output. Either case, if the features are good enough, people will use them, but none of them will be able to reproduce the features. Profit.</p>\n\n<p>I don't know, maybe I'm too cynical.</p>",
      "rawMarkdown": "I agree it might not be a huge problem in this case. It would still be unfair, if someone spent money on bandwidth and hardware to download the data and calculate the features early on, and in the last couple weeks they are just dumped here.\r\n\r\nWhat is more concerning imho, is if this becomes commonplace. In the very least, it escalates the current problems with scripts to a higher level.\r\n\r\nAnd to add to that, I could use commercial software to calculate some decisive features and post them here. Or I could use open source tools, but manually tweak 5% of the output. Either case, if the features are good enough, people will use them, but none of them will be able to reproduce the features. Profit.\r\n\r\nI don't know, maybe I'm too cynical.",
      "votes": null
    },
    {
      "id": "125887",
      "postDate": "07/04/2016 02:38:04",
      "content": "<p>I totally agree that blindly using features posted in the forum is potentially risky. </p>",
      "rawMarkdown": "I totally agree that blindly using features posted in the forum is potentially risky.",
      "votes": null
    },
    {
      "id": "126036",
      "postDate": "07/05/2016 18:06:52",
      "content": "<p>I'm also concerned about this type of situation. Practically though, this won't impact anyone other than a prize winner, in case they were dumb enough to just plain use these features w/o questioning how these were created. Obviously it will have a different type on impact on the whole LB, just because, like scripts, this type of data will bias individual performances (but this is another problem and, honestly, there aren't many left who seem to care about this).\nAs a 2 time Kaggle winner (hey barisumog!) I know there is a big effort the last few days of a competition just to make sure you can reproduce your submissions. If you are up there fighting for a prize, it would be beyond dumb to not double and triple check the reproducibility of results. And if you can't, you just don't pick that submission. I'm sure our friends in the top 10 are paying close attention to it :-)</p>",
      "rawMarkdown": "I'm also concerned about this type of situation. Practically though, this won't impact anyone other than a prize winner, in case they were dumb enough to just plain use these features w/o questioning how these were created. Obviously it will have a different type on impact on the whole LB, just because, like scripts, this type of data will bias individual performances (but this is another problem and, honestly, there aren't many left who seem to care about this).\r\nAs a 2 time Kaggle winner (hey barisumog!) I know there is a big effort the last few days of a competition just to make sure you can reproduce your submissions. If you are up there fighting for a prize, it would be beyond dumb to not double and triple check the reproducibility of results. And if you can't, you just don't pick that submission. I'm sure our friends in the top 10 are paying close attention to it :-)",
      "votes": null
    },
    {
      "id": "126042",
      "postDate": "07/05/2016 19:28:23",
      "content": "<p>Hey Giulio!</p>\n\n<p>As you said, the feature dumps affect the whole leaderboard, but I didn't want to go into a lengthy discussion in that vein, since again as you said, few seem to care.</p>\n\n<p>The only code that gets inspected is those of the prize eligible teams. That's why I asked for some official opinion about them.</p>\n\n<p>I assume the lack of response means &quot;nothing to comment about&quot;. Which I assume means &quot;use them at your own risk&quot;. Which I assume means &quot;it's okay as long as you don't claim a prize&quot;.</p>\n\n<p>And that makes me a sad panda. I guess we're an endangered species.</p>\n\n<p>Good luck to everyone in the final week!</p>",
      "rawMarkdown": "Hey Giulio!\r\n\r\nAs you said, the feature dumps affect the whole leaderboard, but I didn't want to go into a lengthy discussion in that vein, since again as you said, few seem to care.\r\n\r\nThe only code that gets inspected is those of the prize eligible teams. That's why I asked for some official opinion about them.\r\n\r\nI assume the lack of response means \"nothing to comment about\". Which I assume means \"use them at your own risk\". Which I assume means \"it's okay as long as you don't claim a prize\".\r\n\r\nAnd that makes me a sad panda. I guess we're an endangered species.\r\n\r\nGood luck to everyone in the final week!",
      "votes": null
    },
    {
      "id": "126150",
      "postDate": "07/06/2016 20:39:54",
      "content": "<p>I think it would still be fine to use the precomputed hashes available here and claim a prize, since it would be very easy to reproduce the code afterwards. Anyway, you could just claim that it counts as external data!</p>\n\n<p>Competition organizers are usually much more interested in the approach that you take instead of getting every last bit of AUC (I would not want to be the guy who has to implement a 500-model ensemble in a production system!)</p>\n\n<p>Good luck to everyone ;)</p>",
      "rawMarkdown": "I think it would still be fine to use the precomputed hashes available here and claim a prize, since it would be very easy to reproduce the code afterwards. Anyway, you could just claim that it counts as external data!\r\n\r\nCompetition organizers are usually much more interested in the approach that you take instead of getting every last bit of AUC (I would not want to be the guy who has to implement a 500-model ensemble in a production system!)\r\n\r\nGood luck to everyone ;)",
      "votes": null
    },
    {
      "id": "143049",
      "postDate": "11/06/2016 14:37:20",
      "content": "<p>Thank you for this dicussion, i have the same concerns for Bosch competition, which has a tremendous number of features. Now, every kernel i read shows precomputed features, but it's a very time consumming task to obtain them. Kaggle competition are not about &quot;magic features&quot;, it's more about scientific data mining. If kaggle rules are correctly set, the winner should not use feature dump like that, it's pure external information and it breaks the rules if one cannot provide the method for feature extraction. Don't you think?</p>",
      "rawMarkdown": "Thank you for this dicussion, i have the same concerns for Bosch competition, which has a tremendous number of features. Now, every kernel i read shows precomputed features, but it's a very time consumming task to obtain them. Kaggle competition are not about \"magic features\", it's more about scientific data mining. If kaggle rules are correctly set, the winner should not use feature dump like that, it's pure external information and it breaks the rules if one cannot provide the method for feature extraction. Don't you think?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 125822,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "07/03/2016 01:29:58",
      "content": "<p>What rule do you think is being violated?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125827,
      "author_name": "vikulin",
      "author_url": "",
      "post_date": "07/03/2016 05:00:55",
      "content": "<p>inversion, the problem is that results won't be reproducable.</p>\n\n<p>But if somebody using those precomputed hashes wins the prize, they can always write a script to compute them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125835,
      "author_name": "barisumog",
      "author_url": "",
      "post_date": "07/03/2016 08:15:06",
      "content": "<p>[quote=inversion;125822]</p>\n\n<p>What rule do you think is being violated?</p>\n\n<p>[/quote]</p>\n\n<p>I was thinking of common sense, rather than a specific rule. But yeah, let's say reproducibility.</p>\n\n<p>How is this any different than, say, the first place team dumping the probability outputs of their best model? That doesn't violate any specific rules. Would Kaggle / Avito / everyone be okay with using those probs as features?</p>\n\n<p>That being said, it could just be me. I still see these prize competitions as competitive events for a prize, rather than let's-all-get-together-and-see-what-we-come-up-with.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125867,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "07/03/2016 20:08:47",
      "content": "<p>@barisumog - I do see where you are coming from, and initially grimaced at the request as well as the fulfillment  of the pre-computed image hashes. The way I figure it, though, is that these are relatively straight forward calculations, and thus don't necessarily provide any real competitive advantage, especially since scripts have already been posted showing how to do it. In my opinion, this is not much different than posting the image height and width.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125876,
      "author_name": "barisumog",
      "author_url": "",
      "post_date": "07/03/2016 23:41:07",
      "content": "<p>I agree it might not be a huge problem in this case. It would still be unfair, if someone spent money on bandwidth and hardware to download the data and calculate the features early on, and in the last couple weeks they are just dumped here.</p>\n\n<p>What is more concerning imho, is if this becomes commonplace. In the very least, it escalates the current problems with scripts to a higher level.</p>\n\n<p>And to add to that, I could use commercial software to calculate some decisive features and post them here. Or I could use open source tools, but manually tweak 5% of the output. Either case, if the features are good enough, people will use them, but none of them will be able to reproduce the features. Profit.</p>\n\n<p>I don't know, maybe I'm too cynical.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125887,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "07/04/2016 02:38:04",
      "content": "<p>I totally agree that blindly using features posted in the forum is potentially risky. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126036,
      "author_name": "adjgiulio",
      "author_url": "",
      "post_date": "07/05/2016 18:06:52",
      "content": "<p>I'm also concerned about this type of situation. Practically though, this won't impact anyone other than a prize winner, in case they were dumb enough to just plain use these features w/o questioning how these were created. Obviously it will have a different type on impact on the whole LB, just because, like scripts, this type of data will bias individual performances (but this is another problem and, honestly, there aren't many left who seem to care about this).\nAs a 2 time Kaggle winner (hey barisumog!) I know there is a big effort the last few days of a competition just to make sure you can reproduce your submissions. If you are up there fighting for a prize, it would be beyond dumb to not double and triple check the reproducibility of results. And if you can't, you just don't pick that submission. I'm sure our friends in the top 10 are paying close attention to it :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126042,
      "author_name": "barisumog",
      "author_url": "",
      "post_date": "07/05/2016 19:28:23",
      "content": "<p>Hey Giulio!</p>\n\n<p>As you said, the feature dumps affect the whole leaderboard, but I didn't want to go into a lengthy discussion in that vein, since again as you said, few seem to care.</p>\n\n<p>The only code that gets inspected is those of the prize eligible teams. That's why I asked for some official opinion about them.</p>\n\n<p>I assume the lack of response means &quot;nothing to comment about&quot;. Which I assume means &quot;use them at your own risk&quot;. Which I assume means &quot;it's okay as long as you don't claim a prize&quot;.</p>\n\n<p>And that makes me a sad panda. I guess we're an endangered species.</p>\n\n<p>Good luck to everyone in the final week!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126150,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "07/06/2016 20:39:54",
      "content": "<p>I think it would still be fine to use the precomputed hashes available here and claim a prize, since it would be very easy to reproduce the code afterwards. Anyway, you could just claim that it counts as external data!</p>\n\n<p>Competition organizers are usually much more interested in the approach that you take instead of getting every last bit of AUC (I would not want to be the guy who has to implement a 500-model ensemble in a production system!)</p>\n\n<p>Good luck to everyone ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 143049,
      "author_name": "pyd1data",
      "author_url": "",
      "post_date": "11/06/2016 14:37:20",
      "content": "<p>Thank you for this dicussion, i have the same concerns for Bosch competition, which has a tremendous number of features. Now, every kernel i read shows precomputed features, but it's a very time consumming task to obtain them. Kaggle competition are not about &quot;magic features&quot;, it's more about scientific data mining. If kaggle rules are correctly set, the winner should not use feature dump like that, it's pure external information and it breaks the rules if one cannot provide the method for feature extraction. Don't you think?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "125809": "There are multiple precomputed image hash dumps on the forums. What will Kaggle's / Avito's stance be, if a prize eligible team were using these dumps as features?",
    "125822": "What rule do you think is being violated?",
    "125827": "inversion, the problem is that results won't be reproducable.\r\n\r\nBut if somebody using those precomputed hashes wins the prize, they can always write a script to compute them.",
    "125835": "[quote=inversion;125822]\r\n\r\nWhat rule do you think is being violated?\r\n\r\n[/quote]\r\n\r\nI was thinking of common sense, rather than a specific rule. But yeah, let's say reproducibility.\r\n\r\nHow is this any different than, say, the first place team dumping the probability outputs of their best model? That doesn't violate any specific rules. Would Kaggle / Avito / everyone be okay with using those probs as features?\r\n\r\nThat being said, it could just be me. I still see these prize competitions as competitive events for a prize, rather than let's-all-get-together-and-see-what-we-come-up-with.",
    "125867": "barisumog - I do see where you are coming from, and initially grimaced at the request as well as the fulfillment  of the pre-computed image hashes. The way I figure it, though, is that these are relatively straight forward calculations, and thus don't necessarily provide any real competitive advantage, especially since scripts have already been posted showing how to do it. In my opinion, this is not much different than posting the image height and width.",
    "125876": "I agree it might not be a huge problem in this case. It would still be unfair, if someone spent money on bandwidth and hardware to download the data and calculate the features early on, and in the last couple weeks they are just dumped here.\r\n\r\nWhat is more concerning imho, is if this becomes commonplace. In the very least, it escalates the current problems with scripts to a higher level.\r\n\r\nAnd to add to that, I could use commercial software to calculate some decisive features and post them here. Or I could use open source tools, but manually tweak 5% of the output. Either case, if the features are good enough, people will use them, but none of them will be able to reproduce the features. Profit.\r\n\r\nI don't know, maybe I'm too cynical.",
    "125887": "I totally agree that blindly using features posted in the forum is potentially risky.",
    "126036": "I'm also concerned about this type of situation. Practically though, this won't impact anyone other than a prize winner, in case they were dumb enough to just plain use these features w/o questioning how these were created. Obviously it will have a different type on impact on the whole LB, just because, like scripts, this type of data will bias individual performances (but this is another problem and, honestly, there aren't many left who seem to care about this).\r\nAs a 2 time Kaggle winner (hey barisumog!) I know there is a big effort the last few days of a competition just to make sure you can reproduce your submissions. If you are up there fighting for a prize, it would be beyond dumb to not double and triple check the reproducibility of results. And if you can't, you just don't pick that submission. I'm sure our friends in the top 10 are paying close attention to it :-)",
    "126042": "Hey Giulio!\r\n\r\nAs you said, the feature dumps affect the whole leaderboard, but I didn't want to go into a lengthy discussion in that vein, since again as you said, few seem to care.\r\n\r\nThe only code that gets inspected is those of the prize eligible teams. That's why I asked for some official opinion about them.\r\n\r\nI assume the lack of response means \"nothing to comment about\". Which I assume means \"use them at your own risk\". Which I assume means \"it's okay as long as you don't claim a prize\".\r\n\r\nAnd that makes me a sad panda. I guess we're an endangered species.\r\n\r\nGood luck to everyone in the final week!",
    "126150": "I think it would still be fine to use the precomputed hashes available here and claim a prize, since it would be very easy to reproduce the code afterwards. Anyway, you could just claim that it counts as external data!\r\n\r\nCompetition organizers are usually much more interested in the approach that you take instead of getting every last bit of AUC (I would not want to be the guy who has to implement a 500-model ensemble in a production system!)\r\n\r\nGood luck to everyone ;)",
    "143049": "Thank you for this dicussion, i have the same concerns for Bosch competition, which has a tremendous number of features. Now, every kernel i read shows precomputed features, but it's a very time consumming task to obtain them. Kaggle competition are not about \"magic features\", it's more about scientific data mining. If kaggle rules are correctly set, the winner should not use feature dump like that, it's pure external information and it breaks the rules if one cannot provide the method for feature extraction. Don't you think?"
  },
  "source": "meta"
}