{"cells":[
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "#Latent Destination Features\n##Let's have a look at the latent search region features.\n##It won't help you to boost your score immediately although you might gain a few ideas how to apply dimensionality reduction.\n\n##I just wanted to play with seaborn a bit.\n"
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": ""
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "%matplotlib inline\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set_style('whitegrid')\nsns.set(color_codes=True)"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "Destinations.csv has 62K rows. Let's keep only the frequent search destinations. \nRemoving 50K records we could still keep 97% of the test bookings."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "destination_features = pd.read_csv(\"../input/destinations.csv\")\nprint(destination_features.shape)\ntest_destinations = pd.read_csv(\"../input/test.csv\", usecols=['srch_destination_id'])\nsrch_destinations, count = np.unique(test_destinations, return_counts=True)\nfig, ax = plt.subplots(ncols=2, sharex=True)\nax[0].semilogy(sorted(count))\nax[1].plot(1.0 * np.array(sorted(count)).cumsum()/count.sum())\nax[0].set_xticks(range(0, len(srch_destinations), 10000))\nax[1].set_ylabel('Cumulative sum')\nax[0].set_ylabel('Search destination counts in test set (log scale)')\nfrequent_destinations = srch_destinations[count >= 10]\nprint (1. * count[count >= 10].sum() / count.sum())"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "We show the correlations among the 149 latent features."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "frequent_destinations = srch_destinations[count >= 10]\nfrequent_destination_features = destination_features[destination_features['srch_destination_id'].isin(frequent_destinations)]\nfrequent_destination_features = frequent_destination_features.drop('srch_destination_id', axis=1)\nprint(frequent_destination_features.shape)\ncorrelations = frequent_destination_features.corr()\nf = plt.figure()\nax = sns.heatmap(correlations)\nax.set_xticks([])\nax.set_yticks([])\nplt.title('Tartan or correlation matrix')\nf.savefig('tartan.png', dpi=300)\nplt.show()"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "It looks like a nice tartan! It is easy to see that we have many strong correlations and the column order seems to be randomized.\n"
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "fig=plt.figure()\nsns.distplot(correlations.values.reshape(correlations.size), bins=50, color='g')\nplt.title('Correlation values')\nplt.show()\nfig.savefig('CorrelationHist')"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "Using hierarchical clustering we try to reorder the features."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "g = sns.clustermap(correlations)\ng.ax_heatmap.set_xticks([])\ng.ax_heatmap.set_yticks([])\ng.savefig('clustermap.png', dpi=300)"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "Use dendogram_col.reordered_ind to get the index of the original columns."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "print(g.dendrogram_col.reordered_ind)"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "Select a few features from the beginning and check their distributions."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "a = [8, 102, 120, 127, 74]\nto_plot = frequent_destination_features[frequent_destination_features.columns[a]].sample(1000)\ng = sns.PairGrid(to_plot, size=3)\ng.map_upper(plt.scatter, s=5, alpha=0.3)\ng.map_lower(sns.kdeplot, cmap=\"Blues_d\")\ng.map_diag(sns.kdeplot, legend=False, shade=True)\nplt.suptitle('A few features')\ng.savefig('cluster_1.png', dpi=300)  \n\n"
 },
 {
  "cell_type": "markdown",
  "metadata": {},
  "source": "Select a few correlated features from the middle and check their distributions."
 },
 {
  "cell_type": "code",
  "execution_count": null,
  "metadata": {
   "collapsed": false
  },
  "outputs": [],
  "source": "b = [89, 69, 115, 105, 71]\ndef green_kde_hack(x, color, **kwargs):\n    sns.kdeplot(x, color='g', **kwargs)\nto_plot = frequent_destination_features[frequent_destination_features.columns[b]].sample(1000)\ng = sns.PairGrid(to_plot, size=3)\ng.map_upper(plt.scatter, s=5, alpha=0.3, color='g')\ng.map_lower(sns.kdeplot, cmap=\"Greens_d\")\ng.map_diag(green_kde_hack, legend=False, shade=True)\nplt.suptitle('A few correlated features')\ng.savefig('cluster_2.png', dpi=300)"
 }
],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}}, "nbformat": 4, "nbformat_minor": 0}