{
  "id": 24970,
  "title": "wrapping up!",
  "url": "/competitions/painter-by-numbers/discussion/24970",
  "author_name": "",
  "post_date": "2016-11-01T00:06:18.020Z",
  "votes": 7,
  "comment_count": 6,
  "views": 659,
  "content": "<p>Hi everyone! Thanks for your interest in this challenge and congrats to everyone who put in the time and effort to beat the RandomForestClassifier benchmark! For the three of you who score &gt; 0.9, I think that a lot of people are curious to hear more about your strategies!</p>\n\n<p>In order to discourage cheating, I included some work by obscure living artists whom I'd like to acknowledge now -</p>\n\n<ul>\n<li>Lisa Rasmussen - <a href=\"http://www.rasmussenillustrationanddesign.com/new-index/\">http://www.rasmussenillustrationanddesign.com/new-index/</a></li>\n<li>Andrea Loney - <a href=\"https://anggeanouveau.com/portfolio/\">https://anggeanouveau.com/portfolio/</a></li>\n<li>Garrett Ames-Ledbetter - <a href=\"http://garrettal.com/\">http://garrettal.com/</a></li>\n<li>Kirsty Ellen Mottram - <a href=\"https://www.facebook.com/KirstysSketchbook/\">https://www.facebook.com/KirstysSketchbook/</a></li>\n</ul>\n\n<p>For sport, I also slipped a famous forger into the test set. Han van Meegeren was a Dutch artist who created several paintings in the style of Dutch masters from the 1700's. van Meegeren was bitter about the lack of commercial success he had with his own art, so he took up creating &quot;new&quot; paintings by old masters and he even sold some forgeries to the Nazis during World War II. I generated a pairwise comparison table for orange-nejc's predictions for van Meegeren and Vermeer paintings in the test set:</p>\n\n<p><img src=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/vermeer_vs_van_meegeren.png?raw=true\" alt=\"enter image description here\" title></p>\n\n<p>Excitingly, orange-nejc generally predicts that Vermeer paintings are more similar to each other than van Meegeren's paintings are to each other or van Meegeren's paintings are to Vermeer's paintings. </p>\n\n<p>One of the goals of this competition was to determine if the algorithm could extrapolate to images by artists who didn't have work in the training set. I've computed the AUC for orange-nejc's predictions for two cases:</p>\n\n<ul>\n<li>0.94218 (pairs of images by artists with work in the training data\nset) </li>\n<li>0.82509 (pairs of images by artists with no work in the training data\nset)     </li>\n</ul>\n\n<p>I should do some boot-strapping to generate meaningful error bars, but it looks like orange-nejc's algorithm is better at interpolating from artists whose work it has seen before than at extrapolating to work by unfamiliar artists.</p>\n\n<p>For more colourful plots and discussion of how I set up the competition are available on  <a href=\"http://small-yellow-duck.github.io/painter_by_numbers.html\">my github page</a></p>",
  "messages": [
    {
      "id": "142241",
      "postDate": "11/01/2016 00:06:18",
      "content": "<p>Hi everyone! Thanks for your interest in this challenge and congrats to everyone who put in the time and effort to beat the RandomForestClassifier benchmark! For the three of you who score &gt; 0.9, I think that a lot of people are curious to hear more about your strategies!</p>\n\n<p>In order to discourage cheating, I included some work by obscure living artists whom I'd like to acknowledge now -</p>\n\n<ul>\n<li>Lisa Rasmussen - <a href=\"http://www.rasmussenillustrationanddesign.com/new-index/\">http://www.rasmussenillustrationanddesign.com/new-index/</a></li>\n<li>Andrea Loney - <a href=\"https://anggeanouveau.com/portfolio/\">https://anggeanouveau.com/portfolio/</a></li>\n<li>Garrett Ames-Ledbetter - <a href=\"http://garrettal.com/\">http://garrettal.com/</a></li>\n<li>Kirsty Ellen Mottram - <a href=\"https://www.facebook.com/KirstysSketchbook/\">https://www.facebook.com/KirstysSketchbook/</a></li>\n</ul>\n\n<p>For sport, I also slipped a famous forger into the test set. Han van Meegeren was a Dutch artist who created several paintings in the style of Dutch masters from the 1700's. van Meegeren was bitter about the lack of commercial success he had with his own art, so he took up creating &quot;new&quot; paintings by old masters and he even sold some forgeries to the Nazis during World War II. I generated a pairwise comparison table for orange-nejc's predictions for van Meegeren and Vermeer paintings in the test set:</p>\n\n<p><img src=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/vermeer_vs_van_meegeren.png?raw=true\" alt=\"enter image description here\" title></p>\n\n<p>Excitingly, orange-nejc generally predicts that Vermeer paintings are more similar to each other than van Meegeren's paintings are to each other or van Meegeren's paintings are to Vermeer's paintings. </p>\n\n<p>One of the goals of this competition was to determine if the algorithm could extrapolate to images by artists who didn't have work in the training set. I've computed the AUC for orange-nejc's predictions for two cases:</p>\n\n<ul>\n<li>0.94218 (pairs of images by artists with work in the training data\nset) </li>\n<li>0.82509 (pairs of images by artists with no work in the training data\nset)     </li>\n</ul>\n\n<p>I should do some boot-strapping to generate meaningful error bars, but it looks like orange-nejc's algorithm is better at interpolating from artists whose work it has seen before than at extrapolating to work by unfamiliar artists.</p>\n\n<p>For more colourful plots and discussion of how I set up the competition are available on  <a href=\"http://small-yellow-duck.github.io/painter_by_numbers.html\">my github page</a></p>",
      "rawMarkdown": "Hi everyone! Thanks for your interest in this challenge and congrats to everyone who put in the time and effort to beat the RandomForestClassifier benchmark! For the three of you who score > 0.9, I think that a lot of people are curious to hear more about your strategies!\r\n\r\nIn order to discourage cheating, I included some work by obscure living artists whom I'd like to acknowledge now -\r\n\r\n - Lisa Rasmussen - http://www.rasmussenillustrationanddesign.com/new-index/\r\n - Andrea Loney - https://anggeanouveau.com/portfolio/\r\n - Garrett Ames-Ledbetter - http://garrettal.com/\r\n - Kirsty Ellen Mottram - https://www.facebook.com/KirstysSketchbook/\r\n\r\nFor sport, I also slipped a famous forger into the test set. Han van Meegeren was a Dutch artist who created several paintings in the style of Dutch masters from the 1700's. van Meegeren was bitter about the lack of commercial success he had with his own art, so he took up creating \"new\" paintings by old masters and he even sold some forgeries to the Nazis during World War II. I generated a pairwise comparison table for orange-nejc's predictions for van Meegeren and Vermeer paintings in the test set:\r\n\r\n![enter image description here][1]\r\n\r\nExcitingly, orange-nejc generally predicts that Vermeer paintings are more similar to each other than van Meegeren's paintings are to each other or van Meegeren's paintings are to Vermeer's paintings. \r\n\r\nOne of the goals of this competition was to determine if the algorithm could extrapolate to images by artists who didn't have work in the training set. I've computed the AUC for orange-nejc's predictions for two cases:\r\n\r\n - 0.94218 (pairs of images by artists with work in the training data\r\n   set)\t\r\n - 0.82509 (pairs of images by artists with no work in the training data\r\n   set)\t\t\r\n\r\nI should do some boot-strapping to generate meaningful error bars, but it looks like orange-nejc's algorithm is better at interpolating from artists whose work it has seen before than at extrapolating to work by unfamiliar artists.\r\n\r\nFor more colourful plots and discussion of how I set up the competition are available on  [my github page][2]\r\n\r\n\r\n  [1]: https://github.com/small-yellow-duck/kaggle_art/blob/master/vermeer_vs_van_meegeren.png?raw=true\r\n  [2]: http://small-yellow-duck.github.io/painter_by_numbers.html",
      "votes": null
    },
    {
      "id": "142242",
      "postDate": "11/01/2016 00:16:33",
      "content": "<p>There have been a couple inquiries about what I'm going to do with the solutions for the test set labels now that the competition is over. Would people like me to make them public?</p>",
      "rawMarkdown": "There have been a couple inquiries about what I'm going to do with the solutions for the test set labels now that the competition is over. Would people like me to make them public?",
      "votes": null
    },
    {
      "id": "142536",
      "postDate": "11/02/2016 21:13:26",
      "content": "<p>Below is a link to the Github repository that contains my solution for the competition alongside a brief description of the approaches I've used and plans for future work.</p>\n\n<p><a href=\"https://github.com/inejc/painters\">https://github.com/inejc/painters</a></p>",
      "rawMarkdown": "Below is a link to the Github repository that contains my solution for the competition alongside a brief description of the approaches I've used and plans for future work.\r\n\r\nhttps://github.com/inejc/painters",
      "votes": null
    },
    {
      "id": "142556",
      "postDate": "11/02/2016 23:41:44",
      "content": "<p>Thanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say &quot;it's a little bit of Rembrandt and a little bit of Vermeer&quot;.</p>",
      "rawMarkdown": "Thanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say \"it's a little bit of Rembrandt and a little bit of Vermeer\".",
      "votes": null
    },
    {
      "id": "142589",
      "postDate": "11/03/2016 07:24:47",
      "content": "<p>Kiri, thanks a lot for a top-quality competition!</p>\n\n<p>4th place solution is probably not of an interest for most, but dropping it along with brief reasoning behind it here anyways :)</p>\n\n<p><a href=\"https://github.com/myurasov/Painters-by-Numbers-v2\">https://github.com/myurasov/Painters-by-Numbers-v2</a></p>\n\n<p>Nejc &#8211; good job taking a hard path and making it perform!</p>",
      "rawMarkdown": "Kiri, thanks a lot for a top-quality competition!\r\n\r\n4th place solution is probably not of an interest for most, but dropping it along with brief reasoning behind it here anyways :)\r\n\r\nhttps://github.com/myurasov/Painters-by-Numbers-v2\r\n\r\nNejc – good job taking a hard path and making it perform!",
      "votes": null
    },
    {
      "id": "142702",
      "postDate": "11/03/2016 23:13:38",
      "content": "<p>[quote=small yellow duck;142556]</p>\n\n<p>Thanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say &quot;it's a little bit of Rembrandt and a little bit of Vermeer&quot;.</p>\n\n<p>[/quote]</p>\n\n<p>I am curious if that would improve the performance. I will get back to you when I start working on the problem again.</p>",
      "rawMarkdown": "[quote=small yellow duck;142556]\r\n\r\nThanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say \"it's a little bit of Rembrandt and a little bit of Vermeer\".\r\n\r\n[/quote]\r\n\r\nI am curious if that would improve the performance. I will get back to you when I start working on the problem again.",
      "votes": null
    },
    {
      "id": "190501",
      "postDate": "06/07/2017 20:00:25",
      "content": "<p>Have the test labels been published? Thanks.</p>",
      "rawMarkdown": "Have the test labels been published? Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 142242,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "11/01/2016 00:16:33",
      "content": "<p>There have been a couple inquiries about what I'm going to do with the solutions for the test set labels now that the competition is over. Would people like me to make them public?</p>",
      "votes": null,
      "replies": [
        {
          "id": 190501,
          "author_name": "intaka",
          "author_url": "",
          "post_date": "06/07/2017 20:00:25",
          "content": "<p>Have the test labels been published? Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 142536,
      "author_name": "nejc92",
      "author_url": "",
      "post_date": "11/02/2016 21:13:26",
      "content": "<p>Below is a link to the Github repository that contains my solution for the competition alongside a brief description of the approaches I've used and plans for future work.</p>\n\n<p><a href=\"https://github.com/inejc/painters\">https://github.com/inejc/painters</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 142556,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "11/02/2016 23:41:44",
      "content": "<p>Thanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say &quot;it's a little bit of Rembrandt and a little bit of Vermeer&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 142589,
      "author_name": "mikhailyurasov",
      "author_url": "",
      "post_date": "11/03/2016 07:24:47",
      "content": "<p>Kiri, thanks a lot for a top-quality competition!</p>\n\n<p>4th place solution is probably not of an interest for most, but dropping it along with brief reasoning behind it here anyways :)</p>\n\n<p><a href=\"https://github.com/myurasov/Painters-by-Numbers-v2\">https://github.com/myurasov/Painters-by-Numbers-v2</a></p>\n\n<p>Nejc &#8211; good job taking a hard path and making it perform!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 142702,
      "author_name": "nejc92",
      "author_url": "",
      "post_date": "11/03/2016 23:13:38",
      "content": "<p>[quote=small yellow duck;142556]</p>\n\n<p>Thanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say &quot;it's a little bit of Rembrandt and a little bit of Vermeer&quot;.</p>\n\n<p>[/quote]</p>\n\n<p>I am curious if that would improve the performance. I will get back to you when I start working on the problem again.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "142241": "Hi everyone! Thanks for your interest in this challenge and congrats to everyone who put in the time and effort to beat the RandomForestClassifier benchmark! For the three of you who score > 0.9, I think that a lot of people are curious to hear more about your strategies!\r\n\r\nIn order to discourage cheating, I included some work by obscure living artists whom I'd like to acknowledge now -\r\n\r\n - Lisa Rasmussen - http://www.rasmussenillustrationanddesign.com/new-index/\r\n - Andrea Loney - https://anggeanouveau.com/portfolio/\r\n - Garrett Ames-Ledbetter - http://garrettal.com/\r\n - Kirsty Ellen Mottram - https://www.facebook.com/KirstysSketchbook/\r\n\r\nFor sport, I also slipped a famous forger into the test set. Han van Meegeren was a Dutch artist who created several paintings in the style of Dutch masters from the 1700's. van Meegeren was bitter about the lack of commercial success he had with his own art, so he took up creating \"new\" paintings by old masters and he even sold some forgeries to the Nazis during World War II. I generated a pairwise comparison table for orange-nejc's predictions for van Meegeren and Vermeer paintings in the test set:\r\n\r\n![enter image description here][1]\r\n\r\nExcitingly, orange-nejc generally predicts that Vermeer paintings are more similar to each other than van Meegeren's paintings are to each other or van Meegeren's paintings are to Vermeer's paintings. \r\n\r\nOne of the goals of this competition was to determine if the algorithm could extrapolate to images by artists who didn't have work in the training set. I've computed the AUC for orange-nejc's predictions for two cases:\r\n\r\n - 0.94218 (pairs of images by artists with work in the training data\r\n   set)\t\r\n - 0.82509 (pairs of images by artists with no work in the training data\r\n   set)\t\t\r\n\r\nI should do some boot-strapping to generate meaningful error bars, but it looks like orange-nejc's algorithm is better at interpolating from artists whose work it has seen before than at extrapolating to work by unfamiliar artists.\r\n\r\nFor more colourful plots and discussion of how I set up the competition are available on  [my github page][2]\r\n\r\n\r\n  [1]: https://github.com/small-yellow-duck/kaggle_art/blob/master/vermeer_vs_van_meegeren.png?raw=true\r\n  [2]: http://small-yellow-duck.github.io/painter_by_numbers.html",
    "142242": "There have been a couple inquiries about what I'm going to do with the solutions for the test set labels now that the competition is over. Would people like me to make them public?",
    "142536": "Below is a link to the Github repository that contains my solution for the competition alongside a brief description of the approaches I've used and plans for future work.\r\n\r\nhttps://github.com/inejc/painters",
    "142556": "Thanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say \"it's a little bit of Rembrandt and a little bit of Vermeer\".",
    "142589": "Kiri, thanks a lot for a top-quality competition!\r\n\r\n4th place solution is probably not of an interest for most, but dropping it along with brief reasoning behind it here anyways :)\r\n\r\nhttps://github.com/myurasov/Painters-by-Numbers-v2\r\n\r\nNejc – good job taking a hard path and making it perform!",
    "142702": "[quote=small yellow duck;142556]\r\n\r\nThanks for the write-up, inejc! For kicks, I'm curious to see if re-running your code with the 1584-dimension softmax layer stripped off changes how well your code performs. My guess is that removing the softmax layer might improve the performance of the algorithm when it is extrapolating to unfamiliar artists who only have work in the test set - the softmax function is acting to force the algorithm to decide which of the 1584 artists in the training set a work is generated by. Leaving the last output as a dense layer might allow the algorithm to say \"it's a little bit of Rembrandt and a little bit of Vermeer\".\r\n\r\n[/quote]\r\n\r\nI am curious if that would improve the performance. I will get back to you when I start working on the problem again.",
    "190501": "Have the test labels been published? Thanks."
  },
  "source": "meta"
}