{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Manifold Learning And Autoencoders - Now with more Fashion!\n\nAuthor: Alexandru Papiu","metadata":{"_uuid":"5f038994eea67358eba2a3a5ded7a61845cb126a","_cell_guid":"6dfe74fc-d444-e723-4639-18b3e0fff20e"}},{"cell_type":"markdown","source":"Note I originally tried this on the MNIST data set but now I am trying it on the fashion MNIST dataset!\n\nSome updates: \n\n- changed dataset from MNIST to fashion MNIST\n- use a CNN as part of the autoencoder.\n- Tried using a sigmoid activation function for the latent space so that the space if bounded in the unit hypercute\n\nThe MNIST competition is slowly coming to an end so I figured I'd try something slightly different - let's try to see if we can get some intuition about the geometry and topology of the MNIST dataset.","metadata":{"_uuid":"afd0a6e2846ee420593b40e9b4a9f4921c0d9904","_cell_guid":"6f86e2ca-4c94-0460-2c3b-99fb6001a18f"}},{"cell_type":"markdown","source":"### Loading required packages and data:","metadata":{"_uuid":"1ccb0d2eed43b91b9a2cc7b6dad624fad6da4cf3","_cell_guid":"1745b4fe-3fb0-d876-2120-5049a9d7162a"}},{"cell_type":"code","source":"%matplotlib inline\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nfrom sklearn import metrics\nfrom sklearn.neighbors import NearestNeighbors\n\nfrom keras.models import Sequential, Model\nfrom keras.layers import Dense, Dropout, Convolution2D, MaxPooling2D, Flatten, Input\nfrom keras.optimizers import adam\nfrom keras.utils.np_utils import to_categorical\nimport keras\n\n%config InlineBackend.figure_format = 'retina'","metadata":{"_uuid":"a97f15834f540b6b50042b69595e1d2271f5e382","_execution_state":"idle","_cell_guid":"ca8de246-c0bd-52c0-e358-1185a9575b3c","execution":{"iopub.status.busy":"2022-08-12T20:01:11.242871Z","iopub.execute_input":"2022-08-12T20:01:11.243143Z","iopub.status.idle":"2022-08-12T20:01:14.358305Z","shell.execute_reply.started":"2022-08-12T20:01:11.243085Z","shell.execute_reply":"2022-08-12T20:01:14.357520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(\"../input/fashionmnist/fashion-mnist_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-12T20:01:14.359239Z","iopub.execute_input":"2022-08-12T20:01:14.359575Z","iopub.status.idle":"2022-08-12T20:01:20.538697Z","shell.execute_reply.started":"2022-08-12T20:01:14.359524Z","shell.execute_reply":"2022-08-12T20:01:20.537979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train = train.iloc[:,1:].values\nX_train = X_train.reshape(X_train.shape[0], 28, 28) #reshape to rectangular\nX_train = X_train/255 #pixel values are 0 - 255 - this makes puts them in the range 0 - 1\n\ny_train = train[\"label\"].values","metadata":{"_uuid":"091d951a2dce8e600a8cffd0d7b4630127e75f81","_execution_state":"idle","_cell_guid":"ce56722d-f117-e84b-dbd2-1cd3b4c8c19d","execution":{"iopub.status.busy":"2022-08-12T20:01:20.540131Z","iopub.execute_input":"2022-08-12T20:01:20.540423Z","iopub.status.idle":"2022-08-12T20:01:20.890586Z","shell.execute_reply.started":"2022-08-12T20:01:20.540372Z","shell.execute_reply":"2022-08-12T20:01:20.889931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#define a function that allows us to see the digits:\ndef show(img):\n    plt.imshow(img.reshape(28,28), cmap = \"gray\", interpolation = \"none\")","metadata":{"_uuid":"b1effa9e9637e61b174ebca06b850bdfb2837823","_execution_state":"idle","_cell_guid":"99de6496-7669-4384-7400-dd25018c6eef","execution":{"iopub.status.busy":"2022-08-12T20:01:20.895124Z","iopub.execute_input":"2022-08-12T20:01:20.896021Z","iopub.status.idle":"2022-08-12T20:01:20.901187Z","shell.execute_reply.started":"2022-08-12T20:01:20.895962Z","shell.execute_reply":"2022-08-12T20:01:20.900020Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's pick one image to be our default image - just so we have a reference point. We'll call this the **\"purse\" image** - simply because, well it's a purse:","metadata":{"_uuid":"212371d17dd9e2b90702905253cea4e440b7ef75","_cell_guid":"7bd3500e-0e42-fb0d-3f63-c822ec983fc9"}},{"cell_type":"code","source":"img = X_train[234]\nshow(img)","metadata":{"_uuid":"d2660d1f22ecb66851b6a1d23487b03cd69b2d00","_execution_state":"idle","_cell_guid":"9d15656d-80c2-201d-9003-beaee13152c4","execution":{"iopub.status.busy":"2022-08-12T20:01:20.902758Z","iopub.execute_input":"2022-08-12T20:01:20.903270Z","iopub.status.idle":"2022-08-12T20:01:21.231352Z","shell.execute_reply.started":"2022-08-12T20:01:20.903215Z","shell.execute_reply":"2022-08-12T20:01:21.230495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.sample(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T20:01:21.234412Z","iopub.execute_input":"2022-08-12T20:01:21.234698Z","iopub.status.idle":"2022-08-12T20:01:21.285008Z","shell.execute_reply.started":"2022-08-12T20:01:21.234640Z","shell.execute_reply":"2022-08-12T20:01:21.284013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok so our images are in a space in which every pixel is a variable or a feature. Since there are 28*28 = 784 pixels per image we can thing of the images as sitting in $\\mathbb{R}^{784}$ - a real 784 dimensional vector space.\n\nThis space is very high dimensional and most dimensionality reduction techniques try to exploit the assumption that not all of these dimension are needed to distinguish between the labels (or more generally extract features or achieve some learning task).\n\nWhere does this intuition come from? To see this let's generate a point uniformly at random in the unit 784 dimensional hypercube and see what it looks like. We pick points in the hypercube simply because we normalized the pixel intensities. \n\nMaybe a 784-hypercube sounds intimidating but it has a very easy mathematical definition: it's the set composed of all  vectors in $\\mathbb{R}^{784}$ that have all coordinates less than or equal to one: $$\\{x \\in \\mathbb{R}^{784} | |x_i| < 1\\} = [0,1]^{784}$$","metadata":{"_uuid":"5344c5ca7ce137a787236e3d0199704363d12d81","_cell_guid":"343a1e51-0f5f-6538-8c66-fe12860ab485"}},{"cell_type":"code","source":"#generating a random 28 by 28 image:\nrand_img = np.random.randint(0, 255, (28, 28))\nrand_img = rand_img/255.0\n\nshow(rand_img)","metadata":{"_uuid":"4bf41988a3e55dc65bce641f035831252b021595","_cell_guid":"befda8b4-94d8-1e73-2668-22c14ed27033","execution":{"iopub.status.busy":"2022-08-12T20:01:21.286124Z","iopub.execute_input":"2022-08-12T20:01:21.286556Z","iopub.status.idle":"2022-08-12T20:01:21.483776Z","shell.execute_reply.started":"2022-08-12T20:01:21.286508Z","shell.execute_reply":"2022-08-12T20:01:21.482902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Doesn't look like anything to me!\n\nWe can try to sample at random many times but unless we get extremely lucky all we'll get is static and nothing resembling an actual digit. This is good empirical evidence that the meaningful images - in this case images of digits, are clustered in smaller dimensional subsets in the original 784 dimensional pixel space. This is what is called the **manifold hypothesis**. And the promise is that if we better understand the structure of the manifold we will have an easier time building machine learning systems.","metadata":{"_uuid":"23d35b55d873d4082db86edfa336a5c1b0ad9b63","_cell_guid":"d5367fa5-7367-39fd-e0c3-664d495457a3"}},{"cell_type":"markdown","source":"Before we try to see how to figure out the manifold structure let's take a closer look at out space. For example what happens if we start at the point of a digit and start traveling in a random direction? Will we get any meaningful images?","metadata":{"_uuid":"ee6c44a688ebc21e6b52b679d54bff1d3e3472d3","_cell_guid":"c710659e-5521-046f-01b1-3ae17a75a584"}},{"cell_type":"code","source":"rand_direction = np.random.rand(28, 28) ","metadata":{"_uuid":"f30a9a3043bec610a5ddc400a0f02706f4a8256a","_cell_guid":"1febcf07-9339-4e08-aafd-f6bffac763d0","execution":{"iopub.status.busy":"2022-08-12T20:01:21.485248Z","iopub.execute_input":"2022-08-12T20:01:21.485845Z","iopub.status.idle":"2022-08-12T20:01:21.490752Z","shell.execute_reply.started":"2022-08-12T20:01:21.485788Z","shell.execute_reply":"2022-08-12T20:01:21.489940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Moving in a random direction away from an image in the 784 dimensional image space:","metadata":{"_uuid":"b4b5a6385f6de7ba5e24447f1feb38cd923cea58","_cell_guid":"6153ce43-3371-5d0d-2ac3-9590096b974d"}},{"cell_type":"code","source":"for i in range(16):\n    plt.subplot(4,4,i+1)\n    show(img + i/4*rand_direction)    \n    plt.xticks([])\n    plt.yticks([])","metadata":{"_uuid":"26510714bd1a5ab5598a9a4ee2d1b2edc91827a2","_cell_guid":"3e434fd0-7b95-7c63-4daf-dfef0b8014f7","execution":{"iopub.status.busy":"2022-08-12T20:01:21.493255Z","iopub.execute_input":"2022-08-12T20:01:21.494035Z","iopub.status.idle":"2022-08-12T20:01:22.137423Z","shell.execute_reply.started":"2022-08-12T20:01:21.493988Z","shell.execute_reply":"2022-08-12T20:01:22.136748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that as we move away from the purse image, the images we encounter become less and less distinguishable. At first we can still see the eight shape but before we know it we're back in static land.\n\nPerhaps a good analogy here is that of a solar system: the surface of our planets are the manifolds we're interested in, one for each digit. Now say you're on the surface of the earth which is a 2-manifold and you start moving in a random direction (let's assume gravity doesn't exist and you can go through solid objects). If you don't understand the structure of earth you'll quickly find yourself in space or inside the earth. But if you instead move within the local earth (say spherical) coordinates you will stay on the surface and get to see all the cool stuff. \n\nThere are however some differences: first of all we're in a much higher dimensional space and we're not sure how many dimensions we need to capture the structure of the digit subspaces. Secondly these subspaces could be really crazy looking - think for example two donuts entangled in some weird way. You could in fact get from one manifold to another without going into static space as all.","metadata":{"_uuid":"55bb69f1303c53d00b6890c6fc36db4c0705f70d","_cell_guid":"fc0091ec-ced4-158e-0a10-c3f76b679e46"}},{"cell_type":"markdown","source":"### Our images' best friends aka Nearest Neighbors:\n\nAnother thing to do to understand better the structure of the image space is to look at what images are closest to the \"eight\" image using some metric. In this case I'll use the sklearn knn wrap with l_2 distance as the metric on the flattened images. ","metadata":{"_uuid":"512e4bc736b6a4c0c0bd6e5cbdcc501e36f26ba0","_cell_guid":"47a34c03-db94-ce74-3080-5e4cab28faaf"}},{"cell_type":"code","source":"X_flat = X_train.reshape(X_train.shape[0], X_train.shape[1]*X_train.shape[2])\n\nknn = NearestNeighbors(5000)\n\nknn.fit(X_flat[:5000])","metadata":{"_uuid":"0e9c821d65ed0f3f676f21e22300d9dbb067b74a","_cell_guid":"a7aa66a8-6d89-23b4-bf89-70715aad36c5","execution":{"iopub.status.busy":"2022-08-12T20:01:22.138445Z","iopub.execute_input":"2022-08-12T20:01:22.138723Z","iopub.status.idle":"2022-08-12T20:01:22.156156Z","shell.execute_reply.started":"2022-08-12T20:01:22.138678Z","shell.execute_reply":"2022-08-12T20:01:22.155377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"distances, neighbors = knn.kneighbors(img.flatten().reshape(1, -1))\nneighbors = neighbors[0]\ndistances = distances[0]","metadata":{"_uuid":"a91381552c308d8d8b5b6120c9821d85dcbfd510","_cell_guid":"963d393b-f844-368c-11b5-9872df2b8869","execution":{"iopub.status.busy":"2022-08-12T20:01:22.157372Z","iopub.execute_input":"2022-08-12T20:01:22.157670Z","iopub.status.idle":"2022-08-12T20:01:22.190956Z","shell.execute_reply.started":"2022-08-12T20:01:22.157608Z","shell.execute_reply":"2022-08-12T20:01:22.190274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Histogram of L_2 distances from the \"purse\" image:","metadata":{"_uuid":"bd7a3fffb8a1701daebea8bd0cc38e3e5aa18511","_cell_guid":"bccda10a-a75e-5032-bb39-3e4d52bfdb1d"}},{"cell_type":"code","source":"plt.hist(distances[1:])","metadata":{"_uuid":"6a97ca162b879743b6c3d39a771398f553d09027","_cell_guid":"1661677c-c0e7-f014-d5ca-a830689f245c","execution":{"iopub.status.busy":"2022-08-12T20:01:22.192224Z","iopub.execute_input":"2022-08-12T20:01:22.192523Z","iopub.status.idle":"2022-08-12T20:01:22.352148Z","shell.execute_reply.started":"2022-08-12T20:01:22.192463Z","shell.execute_reply":"2022-08-12T20:01:22.351404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distances of the first 5000 images from the \"purse\" image is roughly normally distributed - in fact it's much more well behaved than I expected. At first I though I'd see multiple modes and a higher variance given that we have different classes. (update there is a reason for this and it has to do with how diastances are distribution in high dimensions - this requires it's own exploration).","metadata":{"_uuid":"e89a224a6ab6ee5b8635e37b9abcf04f49273c66","_cell_guid":"e9d04607-8708-9ceb-7f84-7314dd195979"}},{"cell_type":"markdown","source":"### 32 Nearest Neighbors for our \"purse\" image:","metadata":{"_uuid":"46e1292f90b79811860aed194ed034a6f04ea501","_cell_guid":"99749f29-0d04-559e-275c-aa02baba7d28"}},{"cell_type":"code","source":"for digit_num, num in enumerate(neighbors[:36]):\n    plt.subplot(6,6,digit_num+1)\n    grid_data = X_train[num]  # reshape from 1d to 2d pixel array\n    show(grid_data)\n    plt.xticks([])\n    plt.yticks([])","metadata":{"_uuid":"bc281817426dd0109557be4e1a09e58764e940ef","_cell_guid":"26ea39ae-3152-0fd5-bce3-5470a39dbe58","execution":{"iopub.status.busy":"2022-08-12T20:01:22.353518Z","iopub.execute_input":"2022-08-12T20:01:22.354076Z","iopub.status.idle":"2022-08-12T20:01:23.751316Z","shell.execute_reply.started":"2022-08-12T20:01:22.354018Z","shell.execute_reply":"2022-08-12T20:01:23.750493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interesting stuff - most of the neighbors are also purses! All in all it looks like KNN would be a decent way to attack this problem - maybe with 5 or 10 neighbors.","metadata":{"_uuid":"452c986c7af14441e62c6a9b248804adef0869a2","_cell_guid":"87d42bad-134e-29c7-baac-632b4484b6bf"}},{"cell_type":"markdown","source":"### Learning the manifold with an autoencoder:\n\nOk so how do we figure out what (combinations of) dimensions in the image space are important? One option would be to hand engineer features - for examples the mean of all pixels is probably a good feature to have. Other worthwhile features would be the slant, and the vertical or horizontal symmetry. \n\nBut we want do to machine learning not hand-craft features because we're lazy and machines tend to better capture important features in messy datasets. There are many ways to try to reduce the dimensionality - hereI am going to use an autoencoder. I like autoencoders because they have a nice intuitive appeal and you can train them relatively fast. \n\nAn **autoencoder ** is a feed- forward neural network who tries to learn a lower dimensional representation of our data. It does that by decreasing the number of layers in the middle of the network and then increasing it back to the dimension of the original image. \n\nSince the autoencoder is forced to reconstruct the images from a smaller representation it discards any variation that it doesn't find useful. Here is a great description form the [Deep Learning](http://www.deeplearningbook.org/contents/autoencoders.html) book:\n\n\"The important principle is that the autoencoder can aﬀord to represent only the variations that are needed to reconstruct training examples. If the data generating distribution concentrates near a low-dimensional manifold, this yields representations that implicitly capture a local coordinate system for this manifold: only the variations tangent to the manifold around x need to correspond to changes in h=f(x).\"","metadata":{"_uuid":"bf0aaa35ac0bf47f27c547743d43023dc200e9bd","_cell_guid":"515b7b57-4abf-a9c4-c9af-7f1e7d1adda8"}},{"cell_type":"markdown","source":"Let's build an autoencoder in keras - It will have 3 hidden layers with 64, 3, and 64 units respectively. Our model will compress the image to a 3-dimensional vector and then try to reconstruct it . Note that we don't use the target y at all, instead we use X_flat for both the input and the target i.e. we're doing unsupervised learning.","metadata":{"_uuid":"b2bec03fe26b9cd5182e29aa330ff07122883c69","_cell_guid":"80f17775-979d-3830-220d-7550d9058506"}},{"cell_type":"code","source":"input_img = Input(shape=(28, 28,1))\n\nencoded = Convolution2D(16, 3, 3, input_shape = (28, 28, 1), activation=\"relu\")(input_img)\nencoded = Flatten()(encoded)\n\nencoded = Dense(3, activation='sigmoid')(encoded) \n\ndecoded = Dense(64, activation='sigmoid')(encoded)\ndecoded = Dense(784, activation = 'sigmoid')(decoded)\n\nautoencoder = Model(input=input_img, output=decoded)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T20:01:23.752511Z","iopub.execute_input":"2022-08-12T20:01:23.753010Z","iopub.status.idle":"2022-08-12T20:01:23.838886Z","shell.execute_reply.started":"2022-08-12T20:01:23.752959Z","shell.execute_reply":"2022-08-12T20:01:23.837901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train = X_train.reshape(X_train.shape[0],28,28,1)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T20:01:31.751882Z","iopub.execute_input":"2022-08-12T20:01:31.752250Z","iopub.status.idle":"2022-08-12T20:01:31.758055Z","shell.execute_reply.started":"2022-08-12T20:01:31.752185Z","shell.execute_reply":"2022-08-12T20:01:31.757037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"opt = keras.optimizers.Adam(lr=0.03)\n\nautoencoder.compile(optimizer = opt, loss = \"mse\")\nautoencoder.fit(X_train, X_flat, batch_size = 128,\n                nb_epoch = 20)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T20:01:31.912722Z","iopub.execute_input":"2022-08-12T20:01:31.913077Z","iopub.status.idle":"2022-08-12T20:02:30.891573Z","shell.execute_reply.started":"2022-08-12T20:01:31.913023Z","shell.execute_reply":"2022-08-12T20:02:30.890904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"encoder = Model(input = input_img, output = encoded)\n\n#building the decoder:\nencoded_input = Input(shape=(3,))\nencoded_layer_1 = autoencoder.layers[-2]\nencoded_layer_2 = autoencoder.layers[-1]\n\n\ndecoder = encoded_layer_1(encoded_input)\ndecoder = encoded_layer_2(decoder)\ndecoder = Model(input=encoded_input, output=decoder)","metadata":{"_uuid":"f021e626e1c5eb3ec273d5cd2a936a6b4d758666","_cell_guid":"4f920618-7361-a8f3-57e4-c8bbfaedff3f","execution":{"iopub.status.busy":"2022-08-12T19:39:19.071506Z","iopub.execute_input":"2022-08-12T19:39:19.071833Z","iopub.status.idle":"2022-08-12T19:39:19.090715Z","shell.execute_reply.started":"2022-08-12T19:39:19.071737Z","shell.execute_reply":"2022-08-12T19:39:19.089867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3D - representation learned by the autoencoder:","metadata":{"_uuid":"24ca7794f954bbacac34dcd9ba9ef98f969f8d0b","_cell_guid":"6d39ac90-e759-c4ba-5fc2-e97414942091"}},{"cell_type":"code","source":"import seaborn as sns\n\nX_proj = encoder.predict(X_train[:10000])\nX_proj.shape\n\nproj = pd.DataFrame(X_proj)\nproj.columns = [\"comp_1\", \"comp_2\", \"comp3\"]\nproj[\"labels\"] = y_train[:10000]\n#sns.lmplot(\"comp_1\", \"comp_2\",hue = \"labels\", data = proj, fit_reg=False)\n\nsns.pairplot(data = proj, hue = \"labels\")","metadata":{"_uuid":"fde6d4c8694272391db41983762aa398f1e884d2","_cell_guid":"18294c61-7ee2-4ac1-5ab0-a4e9035c7a30","execution":{"iopub.status.busy":"2022-08-12T19:39:19.092049Z","iopub.execute_input":"2022-08-12T19:39:19.092528Z","iopub.status.idle":"2022-08-12T19:39:24.302677Z","shell.execute_reply.started":"2022-08-12T19:39:19.092478Z","shell.execute_reply":"2022-08-12T19:39:24.301885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#sns.kdeplot(x=\"comp_1\", y=\"comp_2\",hue = \"labels\", data = proj, fit_reg=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T19:39:24.303805Z","iopub.execute_input":"2022-08-12T19:39:24.304233Z","iopub.status.idle":"2022-08-12T19:39:24.308114Z","shell.execute_reply.started":"2022-08-12T19:39:24.304181Z","shell.execute_reply":"2022-08-12T19:39:24.307182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see the autoencoder does a decent job of separating certain classes like 1, 0 and 4. It does better than PCA but is not as good as TSNE. The autoencoder learns a better representation that simpler methods like PCA because it can detect nonlinearities in the data due to its relu activations. In fact if we used linear activation functions and only one hidden layer we would have recovered the PCA case.\n\n\nCan we recover the images from their 2-dimensional represetation?","metadata":{"_uuid":"5d0dee96ab950325c33f03f82e7e13be67943ae8","_cell_guid":"28aa31bf-a848-32ce-16fe-160c97695063"}},{"cell_type":"code","source":"\n# #how well does the autoencoder decode:w1\n# plt.subplot(2,2,1)\n# show(X_train[160])\n# plt.subplot(2,2,2)\n# show(autoencoder.predict(np.expand_dims(X_train[160].flatten(), 0)).reshape(28, 28))\n# plt.subplot(2,2,3)\n# show(X_train[150])\n# plt.subplot(2,2,4)\n# show(autoencoder.predict(np.expand_dims(X_train[150].flatten(), 0)).reshape(28, 28))","metadata":{"_uuid":"cd41ecae450858e31dd41769583b96cdb5663206","_cell_guid":"f85301be-7057-2760-d75d-59102a87fff0","execution":{"iopub.status.busy":"2022-08-12T19:39:24.309747Z","iopub.execute_input":"2022-08-12T19:39:24.310032Z","iopub.status.idle":"2022-08-12T19:39:24.319630Z","shell.execute_reply.started":"2022-08-12T19:39:24.309975Z","shell.execute_reply":"2022-08-12T19:39:24.318813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Not really - the encoding decoding process is quite lossy - but that makes sense since we're converting a 784 dimensional vector into 2 dimensions.","metadata":{"_uuid":"39f0bfe2a85c7c7b5b92de8ab57c792d13f6f2c1","_cell_guid":"30ca9950-d90d-93e9-4117-be04c340a219"}},{"cell_type":"markdown","source":"### Generating new digits by moving in the latent 3D - space:","metadata":{"_uuid":"1bf66d2c608a55193d4084f88a103535b7165464","_cell_guid":"15ee9850-abfb-d609-980f-2457a1f6e211"}},{"cell_type":"markdown","source":"Now the hope is that the new 3-D representation of the data is a good coordinate system for the subspace of the data that is actually meaningful i.e. the digits. One (hand-wavy) way to check this is to see what happens if we sample points in the 3-D representation space and move in various directions - do we get some meaningful change in the decoded image of the path or just noise as we did in the original space?","metadata":{"_uuid":"cacf36c2ec4ad439dd30e405edbb45b20ed9ab0b","_cell_guid":"26f3f20c-6525-9b7b-87d8-ac38886068d5"}},{"cell_type":"code","source":"decoder.summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-12T19:39:24.320986Z","iopub.execute_input":"2022-08-12T19:39:24.321585Z","iopub.status.idle":"2022-08-12T19:39:24.336495Z","shell.execute_reply.started":"2022-08-12T19:39:24.321245Z","shell.execute_reply":"2022-08-12T19:39:24.335586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import matplotlib.pyplot as plt\nplt.rcParams[\"figure.figsize\"] = (20,10)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T19:39:24.338004Z","iopub.execute_input":"2022-08-12T19:39:24.338313Z","iopub.status.idle":"2022-08-12T19:39:24.345058Z","shell.execute_reply.started":"2022-08-12T19:39:24.338250Z","shell.execute_reply":"2022-08-12T19:39:24.344268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#moving along the x axis:\nfor i in range(64):\n    plt.subplot(8,8,i+1)\n    pt = np.array([[i/64, i/64, i/64]])\n    show(decoder.predict(pt).reshape((28, 28)))\n    plt.xticks([])\n    plt.yticks([])\n    \nplt.show()\n\nfor i in range(64):\n    plt.subplot(8,8,i+1)\n    pt = np.array([[1/2, i/64, i/64]])\n    show(decoder.predict(pt).reshape((28, 28)))\n    plt.xticks([])\n    plt.yticks([])\n    \nplt.show()\n\n\nfor i in range(64):\n    plt.subplot(8,8,i+1)\n    pt = np.array([[1-i/64, 1-i/64, 1/3]])\n    show(decoder.predict(pt).reshape((28, 28)))\n    plt.xticks([])\n    plt.yticks([])\n    \nplt.show()\n","metadata":{"_uuid":"4c359f06edf9aa1ea25197872e58cd8cf298ce8e","_cell_guid":"7700fd66-b715-3efb-4675-7184c015cc99","execution":{"iopub.status.busy":"2022-08-12T19:39:24.346569Z","iopub.execute_input":"2022-08-12T19:39:24.347026Z","iopub.status.idle":"2022-08-12T19:39:33.458031Z","shell.execute_reply.started":"2022-08-12T19:39:24.346856Z","shell.execute_reply":"2022-08-12T19:39:33.457343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Pretty neat! We see that moving in a given direction in the 3D representation corresponds to staying on the fasion manifold in the original space. We never end up in static space. Sure not all the images are exactly a piece of clothing but they are all clothes-like. You can also clearly see the transition from one class to another.\n\nNote that the way we're generating images here doesn't have very good statistical properties since we have to look at the 2-d plot first. Using a variational autoencoder makes the generative process more rigorous but we'll settle for this.","metadata":{"_uuid":"84306582ec67d75d7698ac4797e1e58df5df5ec2","_cell_guid":"0eaf5d23-9e08-0af7-e7c2-73d6a6fcadad"}},{"cell_type":"markdown","source":"### References:\n\n-  [Autoencoder in Keras](https://blog.keras.io/building-autoencoders-in-keras.html) by Francois Chollet\n\n- [Deep Learning Book Ch 14](http://www.deeplearningbook.org/contents/autoencoders.html) by Ian Goodfellow and Yoshua Bengio and Aaron Courville.","metadata":{"_uuid":"08bde3ae3c37bacce770c383bb59bb3107addc18","_cell_guid":"8a08f815-9a1a-0b5d-8cce-a55472edfa1b"}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}