{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Deep learning model is all about math.So, if you are new to deep learning mathematical knowledge is required. Frameworks(pytorch and tensorflow) is later choice. If you know basic knowloedfe about deep learning in mathematic way, then implementing model with coding will be easy for you.\n\nNote: I won't use math equations. I just explain basic concepts knowledge. Hope this kernel will be helpful to you.\n\n1.Deep learning model  is a function that takes inputs(pixel values)\n\n      Lets say we have an image with 28 by 28 pixels total is 784 pixels and we have 784 inputs.\n      Here is the equation\n                          x1 , x2 ,..., x784 \n      x1 = pixel1 and x784 = pixel784.\n      \n      x is a variable and it holds the pixel value from an image\n      \n      Pixel values is grayscale values ranging from 0 to 255\n      \n      So x1 must be somewhere between o and 255 and so for x2, ...,x784\n      \n2.calculate them (using some variables called weights and bias) and applied to activation function called sigmoid or relu\n      \n      We put use some extra variables weight(w) and bias(x) \n      \n      x1 w1 + x2 2 +.., X784 w784 = z(output)\n      \n      So if we calculate this equation (this is a single image with 784 pixels) output(z) will be large \n      coz inputs (x)values are between 0-255 and w, b are randomly choose(let say w = 5 and b = 2).\n      \n      To minimize output(z), we use sigmoid function sigmoid(z). Sigmiod function produces outputs between 0 and 1, \n      so our output is smaller and computation is much easier\n      \n3.produce predicted outputs(lables)\n      \n      output(predicted) = sigmoid(z) \n      \n4.And then we use loss to measure how correct our prediction is.\n                                        \n      loss = outputs(predicted from the model)-outputs(actual target)\n      \n      loss is also called error. If if the difference is small prediciton will be good but if the difference \n      is large,its a bad prediction\n      \n5.Costfunction is summation of loss\n      \n      cost = sum of losses\n      \n      As we say loss and cost are directly proportional. If loss small, cost will be small and prediction will be \n      near to the desired output.\nAll the steps above are call forward propagation.\n \nDo you remember we put weight and bias put to some value randomly. So prediction(accuracy) will be pretty aweful.And here, backpropagation starts. \n\nThe main goal in deep learning model is to minimize the cost function to nearly zero.\n\n    Backpropagation is calculating inputs to get a desire output. So its a reverse calculating. We dont calculate \n    to get output. We caculate to get input for a desired ouputs.That's called backward.\n\n    Tominimize cost function we use gradient descents. It means differentiate cost with respect to w and b.\n![](http://miro.medium.com/max/1400/1*KQVi812_aERFRolz_5G3rA.gif)                            \n    Then update weight and bias and calculate output again to see cost is closed to zero.(this step is looping \n    till cost is close to zero) and accuracy is better and better\n\nIts all about basic deep learning. And there is a lot to learn. I just cover all concepts.\n\nCNN is easy to learn, so learn it yourself :).\n\nlets write some code. I'll use pytorch.","metadata":{}},{"cell_type":"code","source":"#import libraries\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport torch\nimport torch.nn as nn\nfrom torch.autograd import Variable\nfrom torch.utils.data import DataLoader\nfrom sklearn.model_selection import train_test_split\nimport seaborn as sns\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:08:43.036098Z","iopub.execute_input":"2022-07-10T09:08:43.036568Z","iopub.status.idle":"2022-07-10T09:08:43.046351Z","shell.execute_reply.started":"2022-07-10T09:08:43.036525Z","shell.execute_reply":"2022-07-10T09:08:43.045233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#prepare dataset\n#load data\ntrain = pd.read_csv(r\"../input/digit-recognizer/train.csv\",dtype=np.float32)\ntest = pd.read_csv(r\"../input/digit-recognizer/test.csv\", dtype=np.float32)\nprint(train.shape)\nprint(test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:08:43.057459Z","iopub.execute_input":"2022-07-10T09:08:43.058211Z","iopub.status.idle":"2022-07-10T09:08:48.334776Z","shell.execute_reply.started":"2022-07-10T09:08:43.058157Z","shell.execute_reply":"2022-07-10T09:08:48.333372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#split features (pixels ) and labels (numbers 0 to 9)\ntargets_numpy = train.label.values\nfeatures_numpy = train.iloc[:,train.columns!=\"label\"].values/255 # normalization \n\n#train test split . size of train is 80% and size of test is 20%\nfeatures_train, features_test, targets_train, targets_test = train_test_split(features_numpy, targets_numpy, test_size=0.2, random_state=42)\n\n# create feature and targets tensor for train set\nFeaturesTrain = torch.from_numpy(features_train)\nTargetsTrain = torch.from_numpy(targets_train).type(torch.LongTensor)\n\n\nFeaturesTest = torch.from_numpy(features_test)\nTargetsTest = torch.from_numpy(targets_test).type(torch.LongTensor)\n\n#batch size, epoch and iterations\nbatch_size = 100\nnum_iterations = 10000\nnum_epochs = num_iterations/(len(features_train)/batch_size)\nnum_epochs = int (num_epochs)\n\n#pytorch train and test set\ntrain = torch.utils.data.TensorDataset(FeaturesTrain, TargetsTrain)\ntest = torch.utils.data.TensorDataset(FeaturesTest, TargetsTest)\n\n#dataloader\ntrain_loader = DataLoader(train, batch_size=batch_size, shuffle=False)\ntest_loader = DataLoader(test, batch_size=batch_size, shuffle = False)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:08:48.336788Z","iopub.execute_input":"2022-07-10T09:08:48.337126Z","iopub.status.idle":"2022-07-10T09:08:48.810628Z","shell.execute_reply.started":"2022-07-10T09:08:48.337096Z","shell.execute_reply":"2022-07-10T09:08:48.809664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create CNN Model\nclass CNNModel(nn.Module):\n    def __init__(self):\n        super(CNNModel, self).__init__()\n        \n        # Convolution 1\n        self.cnn1 = nn.Conv2d(in_channels=1, out_channels=16, kernel_size=5, stride=1, padding=0)\n        self.relu1 = nn.ReLU()\n        \n        # Max pool 1\n        self.maxpool1 = nn.MaxPool2d(kernel_size=2)\n     \n        # Convolution 2\n        self.cnn2 = nn.Conv2d(in_channels=16, out_channels=32, kernel_size=5, stride=1, padding=0)\n        self.relu2 = nn.ReLU()\n        \n        # Max pool 2\n        self.maxpool2 = nn.MaxPool2d(kernel_size=2)\n        \n        # Fully connected 1\n        self.fc1 = nn.Linear(32 * 4 * 4, 10) \n    \n    def forward(self, x):\n        # Convolution 1\n        out = self.cnn1(x)\n        out = self.relu1(out)\n        \n        # Max pool 1\n        out = self.maxpool1(out)\n        \n        # Convolution 2 \n        out = self.cnn2(out)\n        out = self.relu2(out)\n        \n        # Max pool 2 \n        out = self.maxpool2(out)\n        \n        # flatten\n        out = out.view(out.size(0), -1)\n\n        # Linear function (readout)\n        out = self.fc1(out)\n        \n        return out\n# Create CNN\nmodel = CNNModel()\n\n# Cross Entropy Loss \nerror = nn.CrossEntropyLoss()\n\n# SGD Optimizer\nlearning_rate = 0.1\noptimizer = torch.optim.SGD(model.parameters(), lr=learning_rate)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:08:48.811867Z","iopub.execute_input":"2022-07-10T09:08:48.812816Z","iopub.status.idle":"2022-07-10T09:08:48.827272Z","shell.execute_reply.started":"2022-07-10T09:08:48.812776Z","shell.execute_reply":"2022-07-10T09:08:48.825941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CNN model training\ncount = 0\nloss_list = []\niteration_list = []\naccuracy_list = []\nfor epoch in range(num_epochs):\n    for i, (images, labels) in enumerate(train_loader):\n        \n        train = Variable(images.view(100,1,28,28))\n        labels = Variable(labels)\n        \n        # Clear gradients\n        optimizer.zero_grad()\n        \n        # Forward propagation\n        outputs = model(train)\n        \n        # Calculate softmax and ross entropy loss\n        loss = error(outputs, labels)\n        \n        # Calculating gradients\n        loss.backward()\n        \n        # Update parameters\n        optimizer.step()\n        \n        count += 1\n        \n        if count % 50 == 0:\n            # Calculate Accuracy         \n            correct = 0\n            total = 0\n            # Iterate through test dataset\n            for images, labels in test_loader:\n                \n                test = Variable(images.view(100,1,28,28))\n                \n                # Forward propagation\n                outputs = model(test)\n                \n                # Get predictions from the maximum value\n                predicted = torch.max(outputs.data, 1)[1]\n                \n                # Total number of labels\n                total += len(labels)\n                \n                correct += (predicted == labels).sum()\n            \n            accuracy = 100 * correct / float(total)\n            \n            # store loss and iteration\n            loss_list.append(loss.data)\n            iteration_list.append(count)\n            accuracy_list.append(accuracy)\n        if count % 500 == 0:\n            # Print Loss\n            print('Iteration: {}  Loss: {}  Accuracy: {} %'.format(count, loss.data, accuracy))","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:08:48.830078Z","iopub.execute_input":"2022-07-10T09:08:48.831086Z","iopub.status.idle":"2022-07-10T09:14:04.698393Z","shell.execute_reply.started":"2022-07-10T09:08:48.831041Z","shell.execute_reply":"2022-07-10T09:14:04.696910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualization loss \nplt.plot(iteration_list,loss_list)\nplt.xlabel(\"Number of iteration\")\nplt.ylabel(\"Loss\")\nplt.title(\"CNN: Loss vs Number of iteration\")\nplt.show()\n\n# visualization accuracy \nplt.plot(iteration_list,accuracy_list,color = \"red\")\nplt.xlabel(\"Number of iteration\")\nplt.ylabel(\"Accuracy\")\nplt.title(\"CNN: Accuracy vs Number of iteration\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:14:04.700404Z","iopub.execute_input":"2022-07-10T09:14:04.700887Z","iopub.status.idle":"2022-07-10T09:14:05.097349Z","shell.execute_reply.started":"2022-07-10T09:14:04.700834Z","shell.execute_reply":"2022-07-10T09:14:05.096134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"my_submission = pd.DataFrame({'ImageId': labels, 'Label': predicted.numpy()})\n# you could use any filename. We choose submission here\nmy_submission.to_csv('sample_submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T09:15:03.015777Z","iopub.execute_input":"2022-07-10T09:15:03.016218Z","iopub.status.idle":"2022-07-10T09:15:03.025903Z","shell.execute_reply.started":"2022-07-10T09:15:03.016184Z","shell.execute_reply":"2022-07-10T09:15:03.024449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}