{
  "id": 20547,
  "title": "Welcome!",
  "url": "/competitions/painter-by-numbers/discussion/20547",
  "author_name": "",
  "post_date": "2016-04-29T18:17:05.353Z",
  "votes": 8,
  "comment_count": 11,
  "views": 2271,
  "content": "<p>Hello fellow Kagglers!</p>\n\n<p>I've been working industriously to prepare this competition and I hope that a few of you are excited about it as I am! I'm looking forward to learning more about neural networks and hopefully coming up with some improvements for the current tools, especially:</p>\n\n<ul>\n<li>as the images in this data set are all different sizes, I'd like to see a solution which is able to accept input images of different dimensions instead of first resizing all the images so that they have the same width and height. </li>\n<li>I'd also like to see some strategies for identifying the  &quot;interesting&quot; parts of an image for the algorithm to focus on.</li>\n</ul>\n\n<p>Finally, I'm curious to see if the pairwise-comparison formulation allows us to train an algorithm which can extrapolate an understanding of artistic style to images by artists whose work it has never viewed before.</p>\n\n<p>Strategies we may want to consider? (Please feel free to share your own!)</p>\n\n<p>Data augmentation</p>\n\n<ul>\n<li>augment the data set by flipping the images left-to-right and labeling the mirrored images as belonging to a new artist</li>\n<li>augment the data set by making small adjustments to the colour balance in the images</li>\n<li>augment the data set by converting the images to greyscale</li>\n<li>augment the data set by decreasing the resolution of the images</li>\n<li>augment the data set by cropping the images randomly (or developing a scheme to crop the images)</li>\n</ul>\n\n<p>Other strategies that might be interesting</p>\n\n<ul>\n<li>siamese neural networks which examine several images simultaneously</li>\n</ul>\n\n<p>Leakage</p>\n\n<p>I've identified two possible sources of leakage and I'm curious to see how significant they will be as the competition progresses:</p>\n\n<ul>\n<li>Some artists sign their paintings.</li>\n<li>Works by a single artist may have been scanned or photographed at the\nsame resolution.</li>\n</ul>\n\n<p>In addition, I've also wondered if the works in the data set are truly high enough resolution - I certainly anticipate that algorithms would perform better with higher resolution images. But hey, let's see what we can do with the data that we do have!</p>",
  "messages": [
    {
      "id": "117577",
      "postDate": "04/29/2016 18:17:05",
      "content": "<p>Hello fellow Kagglers!</p>\n\n<p>I've been working industriously to prepare this competition and I hope that a few of you are excited about it as I am! I'm looking forward to learning more about neural networks and hopefully coming up with some improvements for the current tools, especially:</p>\n\n<ul>\n<li>as the images in this data set are all different sizes, I'd like to see a solution which is able to accept input images of different dimensions instead of first resizing all the images so that they have the same width and height. </li>\n<li>I'd also like to see some strategies for identifying the  &quot;interesting&quot; parts of an image for the algorithm to focus on.</li>\n</ul>\n\n<p>Finally, I'm curious to see if the pairwise-comparison formulation allows us to train an algorithm which can extrapolate an understanding of artistic style to images by artists whose work it has never viewed before.</p>\n\n<p>Strategies we may want to consider? (Please feel free to share your own!)</p>\n\n<p>Data augmentation</p>\n\n<ul>\n<li>augment the data set by flipping the images left-to-right and labeling the mirrored images as belonging to a new artist</li>\n<li>augment the data set by making small adjustments to the colour balance in the images</li>\n<li>augment the data set by converting the images to greyscale</li>\n<li>augment the data set by decreasing the resolution of the images</li>\n<li>augment the data set by cropping the images randomly (or developing a scheme to crop the images)</li>\n</ul>\n\n<p>Other strategies that might be interesting</p>\n\n<ul>\n<li>siamese neural networks which examine several images simultaneously</li>\n</ul>\n\n<p>Leakage</p>\n\n<p>I've identified two possible sources of leakage and I'm curious to see how significant they will be as the competition progresses:</p>\n\n<ul>\n<li>Some artists sign their paintings.</li>\n<li>Works by a single artist may have been scanned or photographed at the\nsame resolution.</li>\n</ul>\n\n<p>In addition, I've also wondered if the works in the data set are truly high enough resolution - I certainly anticipate that algorithms would perform better with higher resolution images. But hey, let's see what we can do with the data that we do have!</p>",
      "rawMarkdown": "Hello fellow Kagglers!\r\n\r\nI've been working industriously to prepare this competition and I hope that a few of you are excited about it as I am! I'm looking forward to learning more about neural networks and hopefully coming up with some improvements for the current tools, especially:\r\n\r\n - as the images in this data set are all different sizes, I'd like to see a solution which is able to accept input images of different dimensions instead of first resizing all the images so that they have the same width and height. \r\n - I'd also like to see some strategies for identifying the  \"interesting\" parts of an image for the algorithm to focus on.\r\n\r\nFinally, I'm curious to see if the pairwise-comparison formulation allows us to train an algorithm which can extrapolate an understanding of artistic style to images by artists whose work it has never viewed before.\r\n\r\nStrategies we may want to consider? (Please feel free to share your own!)\r\n\r\nData augmentation\r\n\r\n - augment the data set by flipping the images left-to-right and labeling the mirrored images as belonging to a new artist\r\n - augment the data set by making small adjustments to the colour balance in the images\r\n - augment the data set by converting the images to greyscale\r\n - augment the data set by decreasing the resolution of the images\r\n - augment the data set by cropping the images randomly (or developing a scheme to crop the images)\r\n\r\nOther strategies that might be interesting\r\n\r\n - siamese neural networks which examine several images simultaneously\r\n\r\n\r\nLeakage\r\n\r\nI've identified two possible sources of leakage and I'm curious to see how significant they will be as the competition progresses:\r\n\r\n - Some artists sign their paintings.\r\n - Works by a single artist may have been scanned or photographed at the\r\n   same resolution.\r\n\r\nIn addition, I've also wondered if the works in the data set are truly high enough resolution - I certainly anticipate that algorithms would perform better with higher resolution images. But hey, let's see what we can do with the data that we do have!",
      "votes": null
    },
    {
      "id": "117578",
      "postDate": "04/29/2016 18:17:23",
      "content": "<p>For anyone looking to get started with siamese convolutional neural networks, a Keras script for the MNIST data set is here:</p>\n\n<p><a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py</a></p>\n\n<p>Script which I used to generated the RandomForestClassifier benchmark:</p>\n\n<p><a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/art_rfc.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/art_rfc.py</a></p>\n\n<p>Script which generates an image of artist similarities clustered by style\n<a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/plot_artist_style_overlaps.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/plot_artist_style_overlaps.py</a>\n<img src=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/artists_clustered_by_style_train_set.png?raw=true\" alt title></p>",
      "rawMarkdown": "For anyone looking to get started with siamese convolutional neural networks, a Keras script for the MNIST data set is here:\r\n\r\nhttps://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py\r\n\r\nScript which I used to generated the RandomForestClassifier benchmark:\r\n\r\nhttps://github.com/small-yellow-duck/kaggle_art/blob/master/art_rfc.py\r\n\r\n\r\nScript which generates an image of artist similarities clustered by style\r\nhttps://github.com/small-yellow-duck/kaggle_art/blob/master/plot_artist_style_overlaps.py\r\n![][1]\r\n\r\n\r\n  [1]: https://github.com/small-yellow-duck/kaggle_art/blob/master/artists_clustered_by_style_train_set.png?raw=true",
      "votes": null
    },
    {
      "id": "118822",
      "postDate": "05/05/2016 14:53:29",
      "content": "<p>Hey, are the images are not available in the scripts?</p>",
      "rawMarkdown": "Hey, are the images are not available in the scripts?",
      "votes": null
    },
    {
      "id": "119036",
      "postDate": "05/06/2016 20:18:06",
      "content": "<p>Because the data sets are pretty large, the images aren't available for scripts. However, the info file for the training set (with image title, style and genre information) is available.</p>",
      "rawMarkdown": "Because the data sets are pretty large, the images aren't available for scripts. However, the info file for the training set (with image title, style and genre information) is available.",
      "votes": null
    },
    {
      "id": "121089",
      "postDate": "05/23/2016 18:44:31",
      "content": "<p>Dear small yellow duck,\nI really appreciate your effort on setting up this competition!\nThis is a really interesting field that has been lacking open and public datasets.</p>\n\n<p>I would like to cite one of the most important papers on this topic:\nJohnson Jr, C. Richard, et al. &quot;Image processing for artist identification.&quot; Signal Processing Magazine, IEEE 25.4 (2008): 37-48.\nEven though they work on the identification of one specific artist, it could be a nice starting point for this verification task. However, the images are not publicly available.</p>\n\n<p>Additionally, there is one work that created a public dataset for multi-class identification problem:\nKhan, Fahad Shahbaz, et al. &quot;Painting-91: a large scale database for computational painting categorization.&quot; Machine vision and applications 25.6 (2014): 1385-1397.\nHowever, the images are not in high resolution.</p>\n\n<p>Therefore, I do think you did an amazing job gathering these high resolution images and making them publicly available.</p>\n\n<p>Recently, co-authors and I have worked on the problem of identifying van Gogh's paintings (similar to the first cited paper above).\nOne of the main challenges was to collect the images and create the dataset. We have crawled over 27,000 pages on Wikimedia and ended up with 333 RGB high resolution images. The paper has been accepted at ICIP 2016 and we will release the dataset as soon as the paper is published.</p>\n\n<p>Given that, I would like to make a few remarks:</p>\n\n<ul>\n<li>Some authors (myself included) believe that density normalization is important in this field. This means that digitized paintings should have a standard resolution, in terms of pixels per painted inch. The value 196.3 has been used before. Having different resolutions could introduce an unwanted bias in classifiers, as well as make the task of extracting patterns from brush strokes more difficult, or even uninformative, due to the lack of normalization.</li>\n<li>To the best of my knowledge, this difference in densities have not been explicitly studied before, though some works had such difference, but didn't take them into consideration.</li>\n</ul>\n\n<p>Finally, my objectives with this post are:</p>\n\n<ul>\n<li>People should also analyze their algorithms based on density information. Maybe evaluate performance based on density normalization, as mixed resolutions should degrade results. How much is lost? To which extent can a single algorithm/system/DNN handle this lack of normalization?</li>\n<li>Can we (you and I) work together on creating a larger and public dataset for this problem, taking such information into account? How have you gathered the data from WikiArt? Do they provide a public API?</li>\n</ul>\n\n<p>Thanks and best regards! And I wish the best of luck to all of you.</p>\n\n<p>PS.: To those interested, I'll post the link to our paper as soon as it is published.</p>\n\n<p>=== EDIT ===</p>\n\n<p>The paper is available at IEEE Xplore (<em>free access until October 6, 2016</em>):\n<a href=\"https://dx.doi.org/10.1109/icip.2016.7532335\">https://dx.doi.org/10.1109/icip.2016.7532335</a></p>\n\n<p>The dataset is available at figshare (<em>CC BY 4.0</em>):\n<a href=\"https://dx.doi.org/10.6084/m9.figshare.3370627\">https://dx.doi.org/10.6084/m9.figshare.3370627</a></p>\n\n<p>The source code is available at GitHub (<em>Apache 2.0</em>):\n<a href=\"https://github.com/gfolego/vangogh\">https://github.com/gfolego/vangogh</a></p>\n\n<p>There is also an entry in Kaggle Datasets:\n<a href=\"https://www.kaggle.com/gfolego/vangogh\">https://www.kaggle.com/gfolego/vangogh</a></p>",
      "rawMarkdown": "Dear small yellow duck,\r\nI really appreciate your effort on setting up this competition!\r\nThis is a really interesting field that has been lacking open and public datasets.\r\n\r\nI would like to cite one of the most important papers on this topic:\r\nJohnson Jr, C. Richard, et al. \"Image processing for artist identification.\" Signal Processing Magazine, IEEE 25.4 (2008): 37-48.\r\nEven though they work on the identification of one specific artist, it could be a nice starting point for this verification task. However, the images are not publicly available.\r\n\r\nAdditionally, there is one work that created a public dataset for multi-class identification problem:\r\nKhan, Fahad Shahbaz, et al. \"Painting-91: a large scale database for computational painting categorization.\" Machine vision and applications 25.6 (2014): 1385-1397.\r\nHowever, the images are not in high resolution.\r\n\r\nTherefore, I do think you did an amazing job gathering these high resolution images and making them publicly available.\r\n\r\nRecently, co-authors and I have worked on the problem of identifying van Gogh's paintings (similar to the first cited paper above).\r\nOne of the main challenges was to collect the images and create the dataset. We have crawled over 27,000 pages on Wikimedia and ended up with 333 RGB high resolution images. The paper has been accepted at ICIP 2016 and we will release the dataset as soon as the paper is published.\r\n\r\nGiven that, I would like to make a few remarks:\r\n\r\n- Some authors (myself included) believe that density normalization is important in this field. This means that digitized paintings should have a standard resolution, in terms of pixels per painted inch. The value 196.3 has been used before. Having different resolutions could introduce an unwanted bias in classifiers, as well as make the task of extracting patterns from brush strokes more difficult, or even uninformative, due to the lack of normalization.\r\n- To the best of my knowledge, this difference in densities have not been explicitly studied before, though some works had such difference, but didn't take them into consideration.\r\n\r\nFinally, my objectives with this post are:\r\n\r\n - People should also analyze their algorithms based on density information. Maybe evaluate performance based on density normalization, as mixed resolutions should degrade results. How much is lost? To which extent can a single algorithm/system/DNN handle this lack of normalization?\r\n - Can we (you and I) work together on creating a larger and public dataset for this problem, taking such information into account? How have you gathered the data from WikiArt? Do they provide a public API?\r\n\r\nThanks and best regards! And I wish the best of luck to all of you.\r\n\r\nPS.: To those interested, I'll post the link to our paper as soon as it is published.\r\n\r\n\r\n=== EDIT ===\r\n\r\nThe paper is available at IEEE Xplore (*free access until October 6, 2016*):\r\nhttps://dx.doi.org/10.1109/icip.2016.7532335\r\n\r\nThe dataset is available at figshare (*CC BY 4.0*):\r\nhttps://dx.doi.org/10.6084/m9.figshare.3370627\r\n\r\nThe source code is available at GitHub (*Apache 2.0*):\r\nhttps://github.com/gfolego/vangogh\r\n\r\nThere is also an entry in Kaggle Datasets:\r\nhttps://www.kaggle.com/gfolego/vangogh",
      "votes": null
    },
    {
      "id": "121096",
      "postDate": "05/23/2016 20:06:52",
      "content": "<p>Guilherme, thanks for your thoughtful comments!</p>\n\n<p>I agree that it would be desirable to have information about number of pixels per square inch of the painting. However, the dimensions of the original paintings are not generally available on wikiart or wikipedia. Putting out a call to crowd-source this data for wikiart would probably be the best way to go about getting the dimensions of the physical paintings. I know some museums include the physical dimensions of the items in their collections.</p>\n\n<p>Although having information about pixels/square inch would be valuable, I still think that the data set as it stands is interesting - after all, people can differentiate a Vermeer from a van Gogh without having to know what resolution photographs of the paintings were taken at. </p>\n\n<p>Since I scraped wikiart, the site has built a public API.</p>",
      "rawMarkdown": "Guilherme, thanks for your thoughtful comments!\r\n\r\nI agree that it would be desirable to have information about number of pixels per square inch of the painting. However, the dimensions of the original paintings are not generally available on wikiart or wikipedia. Putting out a call to crowd-source this data for wikiart would probably be the best way to go about getting the dimensions of the physical paintings. I know some museums include the physical dimensions of the items in their collections.\r\n\r\nAlthough having information about pixels/square inch would be valuable, I still think that the data set as it stands is interesting - after all, people can differentiate a Vermeer from a van Gogh without having to know what resolution photographs of the paintings were taken at. \r\n\r\nSince I scraped wikiart, the site has built a public API.",
      "votes": null
    },
    {
      "id": "121188",
      "postDate": "05/24/2016 19:34:47",
      "content": "<p>Hello small yellow duck,</p>\n\n<p>I am trying to understand the posted Siamese CNN, which you mentioned \n(<a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py</a>).\nWhat I do not understand is how you can train this network when you need the labels for the test set?</p>\n\n<p>If the labels for the test set are available - then why do you need to predict?</p>\n\n<p>It would be great if you can clarify my misunderstanding. Thank you.</p>\n\n<hr>\n\n<p>It is in the below mentioned section the part &quot;(y_test == i)&quot;. </p>\n\n<pre><code># create training+test positive and negative pairs\ndigit_indices = [np.where(y_train == i)[0] for i in range(10)]\ntr_pairs, tr_y = create_pairs(X_train, digit_indices)\n\ndigit_indices = [np.where(y_test == i)[0] for i in range(10)]\nte_pairs, te_y = create_pairs(X_test, digit_indices)\n</code></pre>",
      "rawMarkdown": "Hello small yellow duck,\r\n\r\nI am trying to understand the posted Siamese CNN, which you mentioned \r\n(https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py).\r\nWhat I do not understand is how you can train this network when you need the labels for the test set?\r\n\r\nIf the labels for the test set are available - then why do you need to predict?\r\n\r\nIt would be great if you can clarify my misunderstanding. Thank you.\r\n\r\n\r\n----------\r\n\r\nIt is in the below mentioned section the part \"(y_test == i)\". \r\n\r\n    # create training+test positive and negative pairs\r\n    digit_indices = [np.where(y_train == i)[0] for i in range(10)]\r\n    tr_pairs, tr_y = create_pairs(X_train, digit_indices)\r\n    \r\n    digit_indices = [np.where(y_test == i)[0] for i in range(10)]\r\n    te_pairs, te_y = create_pairs(X_test, digit_indices)",
      "votes": null
    },
    {
      "id": "121193",
      "postDate": "05/24/2016 20:13:27",
      "content": "<p>In the MNIST siamese network tutorial the goal is to evaluate how the algorithm performs on the test set. The model is not trained on the test set - but the predictions on the test set are compared to the labels on the test set to determine how the algorithm has performed.</p>\n\n<p>edit to add: the test set is passed as the validation set so you can see how the algorithm is progressing. If you were to include a criteria for early stopping, you would carve out a chunk of the training data to serve as the validation set. Note that model.fit does not train on the data passed as the validation set.</p>",
      "rawMarkdown": "In the MNIST siamese network tutorial the goal is to evaluate how the algorithm performs on the test set. The model is not trained on the test set - but the predictions on the test set are compared to the labels on the test set to determine how the algorithm has performed.\r\n\r\nedit to add: the test set is passed as the validation set so you can see how the algorithm is progressing. If you were to include a criteria for early stopping, you would carve out a chunk of the training data to serve as the validation set. Note that model.fit does not train on the data passed as the validation set.",
      "votes": null
    },
    {
      "id": "121216",
      "postDate": "05/24/2016 23:15:15",
      "content": "<p>Thank you for the answer, but I probably did not find the best way to describe the difficulty.\nI do understand that in the training the test set is used as validation set.</p>\n\n<p>But In the prediction the test set is also used  as the test set te_pairs depend on the labels?</p>\n\n<pre><code>pred = model.predict([te_pairs[:, 0], te_pairs[:, 1]])\n</code></pre>\n\n<p>The generation of the te_pairs is based on the digit_indices and this uses y_test, therefore depending on the labels of the test set ?!?</p>\n\n<pre><code>digit_indices = [np.where(y_test == i)[0] for i in range(10)]\n</code></pre>\n\n<p>Or is there another way how to get the -probably- needed pairs of the test set without referring to the test set labels?</p>",
      "rawMarkdown": "Thank you for the answer, but I probably did not find the best way to describe the difficulty.\r\nI do understand that in the training the test set is used as validation set.\r\n\r\nBut In the prediction the test set is also used  as the test set te_pairs depend on the labels?\r\n\r\n    pred = model.predict([te_pairs[:, 0], te_pairs[:, 1]])\r\n\r\nThe generation of the te_pairs is based on the digit_indices and this uses y_test, therefore depending on the labels of the test set ?!?\r\n\r\n    digit_indices = [np.where(y_test == i)[0] for i in range(10)]\r\n\r\nOr is there another way how to get the -probably- needed pairs of the test set without referring to the test set labels?",
      "votes": null
    },
    {
      "id": "121223",
      "postDate": "05/25/2016 00:31:16",
      "content": "<p>I think a more stylish method for pairing the images in the test set is to write a routine that randomly pairs images. This would generate image pairs such that 10% of the pairs contained the same number. The existing routine randomly generates image pairs such that 50% of the pairs contain images of the same number. </p>\n\n<p>Remember that the function which generates the test data pairs also needs to generate the labels, te_y, for the test data so that we can evaluate the predictions made by the algorithm - ie, we still need to pass y_test into the pair-generating function because we want to get te_y out.</p>\n\n<p>Please let me know if I haven't answered the question that you're asking!</p>",
      "rawMarkdown": "I think a more stylish method for pairing the images in the test set is to write a routine that randomly pairs images. This would generate image pairs such that 10% of the pairs contained the same number. The existing routine randomly generates image pairs such that 50% of the pairs contain images of the same number. \r\n\r\nRemember that the function which generates the test data pairs also needs to generate the labels, te_y, for the test data so that we can evaluate the predictions made by the algorithm - ie, we still need to pass y_test into the pair-generating function because we want to get te_y out.\r\n\r\nPlease let me know if I haven't answered the question that you're asking!",
      "votes": null
    },
    {
      "id": "373550",
      "postDate": "08/21/2018 14:40:27",
      "content": "<p>Hi, small yellow duck.</p>\n\n<p>I enjoyed your talk at the Google ML/AI event yesterday.  (For the benefit of other readers: the talk was largely about the process of setting up this competition.)  I'm sorry I didn't get a chance to talk with you afterward.  BTW I think the reason for the low participation in this competition was not the lack of prize money but the lack of tiers/ranking points.  I wasn't on Kaggle at the time, but in my own mind winning a prize always seems like an unrealistic hope, while the possibility of advancing my Kaggle status is always an issue.  (Right now I'm technically a discussions and kernels master, but I won't feel like a real Kaggle master until I have a competition gold medal.)</p>\n\n<p>I figure when the video is available I will post a link on some Kaggle-related Slack groups that I'm in, and maybe you will want to post it here.</p>",
      "rawMarkdown": "Hi, small yellow duck.\n\nI enjoyed your talk at the Google ML/AI event yesterday.  (For the benefit of other readers: the talk was largely about the process of setting up this competition.)  I'm sorry I didn't get a chance to talk with you afterward.  BTW I think the reason for the low participation in this competition was not the lack of prize money but the lack of tiers/ranking points.  I wasn't on Kaggle at the time, but in my own mind winning a prize always seems like an unrealistic hope, while the possibility of advancing my Kaggle status is always an issue.  (Right now I'm technically a discussions and kernels master, but I won't feel like a real Kaggle master until I have a competition gold medal.)\n\nI figure when the video is available I will post a link on some Kaggle-related Slack groups that I'm in, and maybe you will want to post it here.",
      "votes": null
    },
    {
      "id": "373789",
      "postDate": "08/21/2018 23:07:44",
      "content": "<p>I had to leave the event as soon as it was over, but I'm happy to bump into you here! I agree with your observation that not awarding Kaggle points for the competition probably deterred people from competing - my recollection is that Kaggle decided not to award points because the data was public.</p>",
      "rawMarkdown": "I had to leave the event as soon as it was over, but I'm happy to bump into you here! I agree with your observation that not awarding Kaggle points for the competition probably deterred people from competing - my recollection is that Kaggle decided not to award points because the data was public.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117578,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "04/29/2016 18:17:23",
      "content": "<p>For anyone looking to get started with siamese convolutional neural networks, a Keras script for the MNIST data set is here:</p>\n\n<p><a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py</a></p>\n\n<p>Script which I used to generated the RandomForestClassifier benchmark:</p>\n\n<p><a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/art_rfc.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/art_rfc.py</a></p>\n\n<p>Script which generates an image of artist similarities clustered by style\n<a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/plot_artist_style_overlaps.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/plot_artist_style_overlaps.py</a>\n<img src=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/artists_clustered_by_style_train_set.png?raw=true\" alt title></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118822,
      "author_name": "arjoonn",
      "author_url": "",
      "post_date": "05/05/2016 14:53:29",
      "content": "<p>Hey, are the images are not available in the scripts?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 119036,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "05/06/2016 20:18:06",
      "content": "<p>Because the data sets are pretty large, the images aren't available for scripts. However, the info file for the training set (with image title, style and genre information) is available.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121089,
      "author_name": "gfolego",
      "author_url": "",
      "post_date": "05/23/2016 18:44:31",
      "content": "<p>Dear small yellow duck,\nI really appreciate your effort on setting up this competition!\nThis is a really interesting field that has been lacking open and public datasets.</p>\n\n<p>I would like to cite one of the most important papers on this topic:\nJohnson Jr, C. Richard, et al. &quot;Image processing for artist identification.&quot; Signal Processing Magazine, IEEE 25.4 (2008): 37-48.\nEven though they work on the identification of one specific artist, it could be a nice starting point for this verification task. However, the images are not publicly available.</p>\n\n<p>Additionally, there is one work that created a public dataset for multi-class identification problem:\nKhan, Fahad Shahbaz, et al. &quot;Painting-91: a large scale database for computational painting categorization.&quot; Machine vision and applications 25.6 (2014): 1385-1397.\nHowever, the images are not in high resolution.</p>\n\n<p>Therefore, I do think you did an amazing job gathering these high resolution images and making them publicly available.</p>\n\n<p>Recently, co-authors and I have worked on the problem of identifying van Gogh's paintings (similar to the first cited paper above).\nOne of the main challenges was to collect the images and create the dataset. We have crawled over 27,000 pages on Wikimedia and ended up with 333 RGB high resolution images. The paper has been accepted at ICIP 2016 and we will release the dataset as soon as the paper is published.</p>\n\n<p>Given that, I would like to make a few remarks:</p>\n\n<ul>\n<li>Some authors (myself included) believe that density normalization is important in this field. This means that digitized paintings should have a standard resolution, in terms of pixels per painted inch. The value 196.3 has been used before. Having different resolutions could introduce an unwanted bias in classifiers, as well as make the task of extracting patterns from brush strokes more difficult, or even uninformative, due to the lack of normalization.</li>\n<li>To the best of my knowledge, this difference in densities have not been explicitly studied before, though some works had such difference, but didn't take them into consideration.</li>\n</ul>\n\n<p>Finally, my objectives with this post are:</p>\n\n<ul>\n<li>People should also analyze their algorithms based on density information. Maybe evaluate performance based on density normalization, as mixed resolutions should degrade results. How much is lost? To which extent can a single algorithm/system/DNN handle this lack of normalization?</li>\n<li>Can we (you and I) work together on creating a larger and public dataset for this problem, taking such information into account? How have you gathered the data from WikiArt? Do they provide a public API?</li>\n</ul>\n\n<p>Thanks and best regards! And I wish the best of luck to all of you.</p>\n\n<p>PS.: To those interested, I'll post the link to our paper as soon as it is published.</p>\n\n<p>=== EDIT ===</p>\n\n<p>The paper is available at IEEE Xplore (<em>free access until October 6, 2016</em>):\n<a href=\"https://dx.doi.org/10.1109/icip.2016.7532335\">https://dx.doi.org/10.1109/icip.2016.7532335</a></p>\n\n<p>The dataset is available at figshare (<em>CC BY 4.0</em>):\n<a href=\"https://dx.doi.org/10.6084/m9.figshare.3370627\">https://dx.doi.org/10.6084/m9.figshare.3370627</a></p>\n\n<p>The source code is available at GitHub (<em>Apache 2.0</em>):\n<a href=\"https://github.com/gfolego/vangogh\">https://github.com/gfolego/vangogh</a></p>\n\n<p>There is also an entry in Kaggle Datasets:\n<a href=\"https://www.kaggle.com/gfolego/vangogh\">https://www.kaggle.com/gfolego/vangogh</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121096,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "05/23/2016 20:06:52",
      "content": "<p>Guilherme, thanks for your thoughtful comments!</p>\n\n<p>I agree that it would be desirable to have information about number of pixels per square inch of the painting. However, the dimensions of the original paintings are not generally available on wikiart or wikipedia. Putting out a call to crowd-source this data for wikiart would probably be the best way to go about getting the dimensions of the physical paintings. I know some museums include the physical dimensions of the items in their collections.</p>\n\n<p>Although having information about pixels/square inch would be valuable, I still think that the data set as it stands is interesting - after all, people can differentiate a Vermeer from a van Gogh without having to know what resolution photographs of the paintings were taken at. </p>\n\n<p>Since I scraped wikiart, the site has built a public API.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121188,
      "author_name": "justfor",
      "author_url": "",
      "post_date": "05/24/2016 19:34:47",
      "content": "<p>Hello small yellow duck,</p>\n\n<p>I am trying to understand the posted Siamese CNN, which you mentioned \n(<a href=\"https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py\">https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py</a>).\nWhat I do not understand is how you can train this network when you need the labels for the test set?</p>\n\n<p>If the labels for the test set are available - then why do you need to predict?</p>\n\n<p>It would be great if you can clarify my misunderstanding. Thank you.</p>\n\n<hr>\n\n<p>It is in the below mentioned section the part &quot;(y_test == i)&quot;. </p>\n\n<pre><code># create training+test positive and negative pairs\ndigit_indices = [np.where(y_train == i)[0] for i in range(10)]\ntr_pairs, tr_y = create_pairs(X_train, digit_indices)\n\ndigit_indices = [np.where(y_test == i)[0] for i in range(10)]\nte_pairs, te_y = create_pairs(X_test, digit_indices)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121193,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "05/24/2016 20:13:27",
      "content": "<p>In the MNIST siamese network tutorial the goal is to evaluate how the algorithm performs on the test set. The model is not trained on the test set - but the predictions on the test set are compared to the labels on the test set to determine how the algorithm has performed.</p>\n\n<p>edit to add: the test set is passed as the validation set so you can see how the algorithm is progressing. If you were to include a criteria for early stopping, you would carve out a chunk of the training data to serve as the validation set. Note that model.fit does not train on the data passed as the validation set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121216,
      "author_name": "justfor",
      "author_url": "",
      "post_date": "05/24/2016 23:15:15",
      "content": "<p>Thank you for the answer, but I probably did not find the best way to describe the difficulty.\nI do understand that in the training the test set is used as validation set.</p>\n\n<p>But In the prediction the test set is also used  as the test set te_pairs depend on the labels?</p>\n\n<pre><code>pred = model.predict([te_pairs[:, 0], te_pairs[:, 1]])\n</code></pre>\n\n<p>The generation of the te_pairs is based on the digit_indices and this uses y_test, therefore depending on the labels of the test set ?!?</p>\n\n<pre><code>digit_indices = [np.where(y_test == i)[0] for i in range(10)]\n</code></pre>\n\n<p>Or is there another way how to get the -probably- needed pairs of the test set without referring to the test set labels?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121223,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "05/25/2016 00:31:16",
      "content": "<p>I think a more stylish method for pairing the images in the test set is to write a routine that randomly pairs images. This would generate image pairs such that 10% of the pairs contained the same number. The existing routine randomly generates image pairs such that 50% of the pairs contain images of the same number. </p>\n\n<p>Remember that the function which generates the test data pairs also needs to generate the labels, te_y, for the test data so that we can evaluate the predictions made by the algorithm - ie, we still need to pass y_test into the pair-generating function because we want to get te_y out.</p>\n\n<p>Please let me know if I haven't answered the question that you're asking!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 373550,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "08/21/2018 14:40:27",
      "content": "<p>Hi, small yellow duck.</p>\n\n<p>I enjoyed your talk at the Google ML/AI event yesterday.  (For the benefit of other readers: the talk was largely about the process of setting up this competition.)  I'm sorry I didn't get a chance to talk with you afterward.  BTW I think the reason for the low participation in this competition was not the lack of prize money but the lack of tiers/ranking points.  I wasn't on Kaggle at the time, but in my own mind winning a prize always seems like an unrealistic hope, while the possibility of advancing my Kaggle status is always an issue.  (Right now I'm technically a discussions and kernels master, but I won't feel like a real Kaggle master until I have a competition gold medal.)</p>\n\n<p>I figure when the video is available I will post a link on some Kaggle-related Slack groups that I'm in, and maybe you will want to post it here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 373789,
          "author_name": "smallyellowduck",
          "author_url": "",
          "post_date": "08/21/2018 23:07:44",
          "content": "<p>I had to leave the event as soon as it was over, but I'm happy to bump into you here! I agree with your observation that not awarding Kaggle points for the competition probably deterred people from competing - my recollection is that Kaggle decided not to award points because the data was public.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "117577": "Hello fellow Kagglers!\r\n\r\nI've been working industriously to prepare this competition and I hope that a few of you are excited about it as I am! I'm looking forward to learning more about neural networks and hopefully coming up with some improvements for the current tools, especially:\r\n\r\n - as the images in this data set are all different sizes, I'd like to see a solution which is able to accept input images of different dimensions instead of first resizing all the images so that they have the same width and height. \r\n - I'd also like to see some strategies for identifying the  \"interesting\" parts of an image for the algorithm to focus on.\r\n\r\nFinally, I'm curious to see if the pairwise-comparison formulation allows us to train an algorithm which can extrapolate an understanding of artistic style to images by artists whose work it has never viewed before.\r\n\r\nStrategies we may want to consider? (Please feel free to share your own!)\r\n\r\nData augmentation\r\n\r\n - augment the data set by flipping the images left-to-right and labeling the mirrored images as belonging to a new artist\r\n - augment the data set by making small adjustments to the colour balance in the images\r\n - augment the data set by converting the images to greyscale\r\n - augment the data set by decreasing the resolution of the images\r\n - augment the data set by cropping the images randomly (or developing a scheme to crop the images)\r\n\r\nOther strategies that might be interesting\r\n\r\n - siamese neural networks which examine several images simultaneously\r\n\r\n\r\nLeakage\r\n\r\nI've identified two possible sources of leakage and I'm curious to see how significant they will be as the competition progresses:\r\n\r\n - Some artists sign their paintings.\r\n - Works by a single artist may have been scanned or photographed at the\r\n   same resolution.\r\n\r\nIn addition, I've also wondered if the works in the data set are truly high enough resolution - I certainly anticipate that algorithms would perform better with higher resolution images. But hey, let's see what we can do with the data that we do have!",
    "117578": "For anyone looking to get started with siamese convolutional neural networks, a Keras script for the MNIST data set is here:\r\n\r\nhttps://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py\r\n\r\nScript which I used to generated the RandomForestClassifier benchmark:\r\n\r\nhttps://github.com/small-yellow-duck/kaggle_art/blob/master/art_rfc.py\r\n\r\n\r\nScript which generates an image of artist similarities clustered by style\r\nhttps://github.com/small-yellow-duck/kaggle_art/blob/master/plot_artist_style_overlaps.py\r\n![][1]\r\n\r\n\r\n  [1]: https://github.com/small-yellow-duck/kaggle_art/blob/master/artists_clustered_by_style_train_set.png?raw=true",
    "118822": "Hey, are the images are not available in the scripts?",
    "119036": "Because the data sets are pretty large, the images aren't available for scripts. However, the info file for the training set (with image title, style and genre information) is available.",
    "121089": "Dear small yellow duck,\r\nI really appreciate your effort on setting up this competition!\r\nThis is a really interesting field that has been lacking open and public datasets.\r\n\r\nI would like to cite one of the most important papers on this topic:\r\nJohnson Jr, C. Richard, et al. \"Image processing for artist identification.\" Signal Processing Magazine, IEEE 25.4 (2008): 37-48.\r\nEven though they work on the identification of one specific artist, it could be a nice starting point for this verification task. However, the images are not publicly available.\r\n\r\nAdditionally, there is one work that created a public dataset for multi-class identification problem:\r\nKhan, Fahad Shahbaz, et al. \"Painting-91: a large scale database for computational painting categorization.\" Machine vision and applications 25.6 (2014): 1385-1397.\r\nHowever, the images are not in high resolution.\r\n\r\nTherefore, I do think you did an amazing job gathering these high resolution images and making them publicly available.\r\n\r\nRecently, co-authors and I have worked on the problem of identifying van Gogh's paintings (similar to the first cited paper above).\r\nOne of the main challenges was to collect the images and create the dataset. We have crawled over 27,000 pages on Wikimedia and ended up with 333 RGB high resolution images. The paper has been accepted at ICIP 2016 and we will release the dataset as soon as the paper is published.\r\n\r\nGiven that, I would like to make a few remarks:\r\n\r\n- Some authors (myself included) believe that density normalization is important in this field. This means that digitized paintings should have a standard resolution, in terms of pixels per painted inch. The value 196.3 has been used before. Having different resolutions could introduce an unwanted bias in classifiers, as well as make the task of extracting patterns from brush strokes more difficult, or even uninformative, due to the lack of normalization.\r\n- To the best of my knowledge, this difference in densities have not been explicitly studied before, though some works had such difference, but didn't take them into consideration.\r\n\r\nFinally, my objectives with this post are:\r\n\r\n - People should also analyze their algorithms based on density information. Maybe evaluate performance based on density normalization, as mixed resolutions should degrade results. How much is lost? To which extent can a single algorithm/system/DNN handle this lack of normalization?\r\n - Can we (you and I) work together on creating a larger and public dataset for this problem, taking such information into account? How have you gathered the data from WikiArt? Do they provide a public API?\r\n\r\nThanks and best regards! And I wish the best of luck to all of you.\r\n\r\nPS.: To those interested, I'll post the link to our paper as soon as it is published.\r\n\r\n\r\n=== EDIT ===\r\n\r\nThe paper is available at IEEE Xplore (*free access until October 6, 2016*):\r\nhttps://dx.doi.org/10.1109/icip.2016.7532335\r\n\r\nThe dataset is available at figshare (*CC BY 4.0*):\r\nhttps://dx.doi.org/10.6084/m9.figshare.3370627\r\n\r\nThe source code is available at GitHub (*Apache 2.0*):\r\nhttps://github.com/gfolego/vangogh\r\n\r\nThere is also an entry in Kaggle Datasets:\r\nhttps://www.kaggle.com/gfolego/vangogh",
    "121096": "Guilherme, thanks for your thoughtful comments!\r\n\r\nI agree that it would be desirable to have information about number of pixels per square inch of the painting. However, the dimensions of the original paintings are not generally available on wikiart or wikipedia. Putting out a call to crowd-source this data for wikiart would probably be the best way to go about getting the dimensions of the physical paintings. I know some museums include the physical dimensions of the items in their collections.\r\n\r\nAlthough having information about pixels/square inch would be valuable, I still think that the data set as it stands is interesting - after all, people can differentiate a Vermeer from a van Gogh without having to know what resolution photographs of the paintings were taken at. \r\n\r\nSince I scraped wikiart, the site has built a public API.",
    "121188": "Hello small yellow duck,\r\n\r\nI am trying to understand the posted Siamese CNN, which you mentioned \r\n(https://github.com/small-yellow-duck/kaggle_art/blob/master/mnist_siamese_cnn.py).\r\nWhat I do not understand is how you can train this network when you need the labels for the test set?\r\n\r\nIf the labels for the test set are available - then why do you need to predict?\r\n\r\nIt would be great if you can clarify my misunderstanding. Thank you.\r\n\r\n\r\n----------\r\n\r\nIt is in the below mentioned section the part \"(y_test == i)\". \r\n\r\n    # create training+test positive and negative pairs\r\n    digit_indices = [np.where(y_train == i)[0] for i in range(10)]\r\n    tr_pairs, tr_y = create_pairs(X_train, digit_indices)\r\n    \r\n    digit_indices = [np.where(y_test == i)[0] for i in range(10)]\r\n    te_pairs, te_y = create_pairs(X_test, digit_indices)",
    "121193": "In the MNIST siamese network tutorial the goal is to evaluate how the algorithm performs on the test set. The model is not trained on the test set - but the predictions on the test set are compared to the labels on the test set to determine how the algorithm has performed.\r\n\r\nedit to add: the test set is passed as the validation set so you can see how the algorithm is progressing. If you were to include a criteria for early stopping, you would carve out a chunk of the training data to serve as the validation set. Note that model.fit does not train on the data passed as the validation set.",
    "121216": "Thank you for the answer, but I probably did not find the best way to describe the difficulty.\r\nI do understand that in the training the test set is used as validation set.\r\n\r\nBut In the prediction the test set is also used  as the test set te_pairs depend on the labels?\r\n\r\n    pred = model.predict([te_pairs[:, 0], te_pairs[:, 1]])\r\n\r\nThe generation of the te_pairs is based on the digit_indices and this uses y_test, therefore depending on the labels of the test set ?!?\r\n\r\n    digit_indices = [np.where(y_test == i)[0] for i in range(10)]\r\n\r\nOr is there another way how to get the -probably- needed pairs of the test set without referring to the test set labels?",
    "121223": "I think a more stylish method for pairing the images in the test set is to write a routine that randomly pairs images. This would generate image pairs such that 10% of the pairs contained the same number. The existing routine randomly generates image pairs such that 50% of the pairs contain images of the same number. \r\n\r\nRemember that the function which generates the test data pairs also needs to generate the labels, te_y, for the test data so that we can evaluate the predictions made by the algorithm - ie, we still need to pass y_test into the pair-generating function because we want to get te_y out.\r\n\r\nPlease let me know if I haven't answered the question that you're asking!",
    "373550": "Hi, small yellow duck.\n\nI enjoyed your talk at the Google ML/AI event yesterday.  (For the benefit of other readers: the talk was largely about the process of setting up this competition.)  I'm sorry I didn't get a chance to talk with you afterward.  BTW I think the reason for the low participation in this competition was not the lack of prize money but the lack of tiers/ranking points.  I wasn't on Kaggle at the time, but in my own mind winning a prize always seems like an unrealistic hope, while the possibility of advancing my Kaggle status is always an issue.  (Right now I'm technically a discussions and kernels master, but I won't feel like a real Kaggle master until I have a competition gold medal.)\n\nI figure when the video is available I will post a link on some Kaggle-related Slack groups that I'm in, and maybe you will want to post it here.",
    "373789": "I had to leave the event as soon as it was over, but I'm happy to bump into you here! I agree with your observation that not awarding Kaggle points for the competition probably deterred people from competing - my recollection is that Kaggle decided not to award points because the data was public."
  },
  "source": "meta"
}