{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "94e71fe2-5fdd-afbd-bffa-d3b332c87722"
      },
      "source": [
        "Essential to always understand the nature of data. Goal of this R workbook is to share what I learnt as I \n",
        "start to look into dataset without doing any tangible data processing. I've released also the ggplot code\n",
        "which I found very useful is understanding, and conveying the message."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "e95dadbd-b60a-6815-e548-a7889fae10ea"
      },
      "outputs": [],
      "source": [
        "# This R environment comes with all of CRAN preinstalled, as well as many other helpful packages\n",
        "# The environment is defined by the kaggle/rstats docker image: https://github.com/kaggle/docker-rstats\n",
        "# For example, here's several helpful packages to load in \n",
        "\n",
        "library(ggplot2) # Data visualization\n",
        "library(readr) # CSV file I/O, e.g. the read_csv function\n",
        "library(dplyr)\n",
        "library(tidyr)\n",
        "\n",
        "# Input data files are available in the \"../input/\" directory.\n",
        "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n",
        "\n",
        "system(\"ls ../input\")\n",
        "\n",
        "# Any results you write to the current directory are saved as output."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "0d4c9889-4f7f-fb40-9b0f-0385f1e8e5df"
      },
      "outputs": [],
      "source": [
        "artists <- read.csv(\"../input/train_info.csv\")\n",
        "str(artists)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "0cec8286-0713-981f-bde9-0e88865b4cbe"
      },
      "source": [
        "Lets see how many artists are in the training "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ac0f0940-19fd-0027-4464-7bdf2bcf3f5e"
      },
      "outputs": [],
      "source": [
        "Lets see how many artists works are in the training"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9d1cb194-a0e2-45f2-107d-455eedd23b31"
      },
      "outputs": [],
      "source": [
        "artists  %>% group_by(artist,genre) %>%  summarise(total = n()) %>%  arrange(desc(total,artist)) %>% ggplot(aes(reorder(artist,total),total)) + \n",
        "geom_bar(stat=\"identity\",width = 0.7) + theme(axis.text.x = element_text(angle = 90)) + ylab(\"Number of works\") + \n",
        "xlab(\"Artist\") +\n",
        "ggtitle(\"No of works by Artist\") +  coord_flip() "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "0c97d5ca-c2a4-3623-8a97-19bb3e13f1dc"
      },
      "source": [
        "Above plot basically tells that same artist has dabbled in many of the genres and his signature\n",
        "is kind of all over the place or is spread. A way to think would be that original dna can be traced, \n",
        "question is how ?"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "fdb77a11-47e8-3cdc-083d-e1051eeb2766"
      },
      "source": [
        "So, have to melt it"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "35788e08-179f-075b-6c88-64e6a124ad75"
      },
      "outputs": [],
      "source": [
        "library(reshape2)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "4320a1c2-c9f7-24cb-b257-6ce6962a9916"
      },
      "outputs": [],
      "source": [
        "artistsbyGenre <- artists  %>% group_by(artist,genre) %>%  \n",
        "summarise(total = n()) %>%  arrange(desc(total,artist)) %>% as.data.frame()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "ae52fbfa-d8b9-2221-0151-9619839439cd"
      },
      "outputs": [],
      "source": [
        "artistsbyGenreMelt<- melt(artistsbyGenre, id.vars = c(\"artist\"))\n",
        "artistsbyGenreMelt<- melt(artistsbyGenre)\n",
        "\n",
        "#plot\n",
        "artistsbyGenreMelt %>% arrange(desc(value)) %>% \n",
        "sample_frac(.20) %>% ggplot(aes(reorder(artist,value), value, fill=genre)) + \n",
        "geom_bar(stat=\"identity\",width = 0.7) + theme(axis.text.x = element_text(angle = 90)) + ylab(\"Value\") + \n",
        "xlab(\"Artist\") +\n",
        "ggtitle(\"Spread by genre\") "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "b076d77d-df20-772d-b5b0-9950022ee74a"
      },
      "source": [
        "Maybe one opyion to try is to group by works of an artist, so read all the works by an artist, \n",
        "and create a feaure vector\n",
        "\n",
        "Lets see which artist is more versatile"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a286ff5d-3ec9-5806-d186-8af7d3cbfe55"
      },
      "outputs": [],
      "source": [
        "artistsbyGenre  %>% \n",
        "group_by(artist) %>%  summarise(total = n())  %>%  arrange(desc(total))  %>% head "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "399ee5c8-ba18-736f-7e51-964cef053528"
      },
      "outputs": [],
      "source": [
        "artistsbyGenre  %>% group_by(artist) %>% arrange(artist)  %>%  \n",
        "filter(artist  %in% c('40f86d376acde0d9862ce7493745bdae') ) %>% as.data.frame() %>% nrow()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "3b4e56f7-fb67-b5e8-471c-3c857ebae1c7"
      },
      "source": [
        "So, there exists a disparity i.e whether by the artist who is te most versaile is not the \n",
        "one who has large corpus? Make sense ? \n",
        "\n",
        "If we need to know the volume"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "6241eacd-6652-589d-c59a-16555de3e88d"
      },
      "outputs": [],
      "source": [
        "artistsbyGenre  %>% group_by(artist) %>% arrange(artist) %>%  \n",
        "summarise(totalworks = sum(total)) %>% arrange(desc(totalworks))  %>% as.data.frame()   %>% nrow()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "c7e49ccf-6e41-b7ec-f736-4f6376b40815"
      },
      "source": [
        "So, there are 1584 artists, not that many\n",
        "Lets see, how many works by genre"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "9d7bf7fd-9df7-3545-b6a2-ff5c6d27c4e2"
      },
      "outputs": [],
      "source": [
        "artistsbyGenre  %>% group_by(genre) %>% arrange(genre) %>%  \n",
        "summarise(totalbygenre = sum(total)) %>% arrange(desc(totalbygenre))  %>% as.data.frame() "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "5f15b6f3-5886-43a5-89bc-32c39ceea12e"
      },
      "source": [
        "So, what we need is the traits of an artist \n",
        "So, how about this, so for each artist we create a blueprint\n",
        "Blueprint will be composed of all his genres "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "197d3d59-36da-4184-3df4-5ef89d3c9fce"
      },
      "source": [
        "Total by genre in a plot, portarit , landscape and genre painting are the most\n",
        "infamous or popular in the training "
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "37510027-397a-d4c5-5a2b-273b71660380"
      },
      "outputs": [],
      "source": [
        "artistsbyGenre  %>% group_by(genre) %>% \n",
        "arrange(genre) %>%  summarise(totalbygenre = sum(total)) %>% arrange(desc(totalbygenre))  %>% \n",
        "as.data.frame() %>% \n",
        "ggplot(aes(genre,totalbygenre)) +  geom_bar(stat = \"identity\") + theme(axis.text.x = element_text(angle = 90, hjust = 1)) "
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "3016940e-959f-b10d-2b6f-cae4cf3f2264"
      },
      "source": [
        "Lets see the spread by style in a genre, \n",
        "this gives the genre, and the spread of styles within it, and the count"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "a820ed37-42c3-1d6c-6cb0-ec913e09eaaf"
      },
      "outputs": [],
      "source": [
        "artistsByGenreStyle <- artists  %>% group_by(genre,style) %>% arrange(genre) %>%  summarise(totalbygenre = n()) %>% as.data.frame()\n",
        "\n",
        "artists  %>% group_by(genre,style) %>% \n",
        "arrange(genre) %>%  summarise(totalbygenre = n()) %>% \n",
        "filter(totalbygenre > 100)%>% as.data.frame() %>% \n",
        "ggplot( aes(style, totalbygenre, width=.85)) +   \n",
        "  geom_bar(aes(fill = genre), position = \"dodge\", stat=\"identity\") +  theme(axis.text.x = element_text(angle = 90, hjust = 1)) + coord_flip()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "457fa5c3-fbac-2d62-5f24-2df155f775cf"
      },
      "source": [
        "So, turn our focus to find the impressionist painters ? Maybe, you will top the leaderboard"
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "R",
      "language": "R",
      "name": "ir"
    },
    "language_info": {
      "codemirror_mode": "r",
      "file_extension": ".r",
      "mimetype": "text/x-r-source",
      "name": "R",
      "pygments_lexer": "r",
      "version": "3.3.1"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}