{"cells":[{"metadata":{},"cell_type":"markdown","source":" # About this notebook\n \nThis notebook presents a Knaive model and is a supporting notebook to the my main notebook here, which presents complete detialed explanation and dicussions abobut this competitiom:\nhttps://www.kaggle.com/tanulsingh077/steganalysis-complete-understanding-and-model ---> Original author (TANUL SINGH)\n\nTo understand how I came up with this idea and how everything works please refer above\n\n**<span style=\"color: red;\">If you like my effort and contribution please show token of gratitude by upvoting my kernels</span>**\n# Basic Idea\nYesterday Kaggle launched the steganalysis competition. I found it very interesting as I have never heard of something like this. It instilled the spy fantasies within me. From yesterday I have been reading research papers on approaching this problem using deep learning but before I tried **SRNET** I want to do something of own. After a lot of thinking I came up with this :-<br><br>\n* In this we are suppose to predict whether the test images are hiding some information or not but we dont have labels for the train images we just clean images and same images encoded using different algos, the main point is creating labels then we can approach this as regression problem . I thought Since in steganography in images, any technique involves changes in pixel values, a very knaive way to get labels would be to flatten the RGB images(both encoded and normal) into a vector and then find cosine dissimilarity between the two vectors, since the encoded value contains a hidden information its vector will differ from the main vector and hence we will have a non zero value of cosine dissimilarity.\n* Then we Label the images as: Cover_images as 0 , JMIPOD images as similarity between cover and JMIPOD images , JUNIWARD images as similarity between cover and JUNIWARD images, UERD images as similarity between cover and UERD images\n* After creating our labels , we will stack all the images and labels in a datafame and then approach this as simple regression problem\n\nUpdate : One flaw that I found in the above approach was that labels to be predicted should be between zero and one, whereas our regression problem predicts values of any range and values can also be negative , so I passed out cosine dissmilarity through sigmoid function to get a probability between zero and one"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# PRELIMINARIES\nimport os\nimport skimage.io as sk\nimport matplotlib.pyplot as plt\nfrom scipy import spatial\nfrom tqdm import tqdm\nfrom PIL import Image\nfrom random import shuffle\n\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"BASE_PATH = \"/kaggle/input/alaska2-image-steganalysis\"\ntrain_imageids = pd.Series(os.listdir(BASE_PATH + '/Cover')).sort_values(ascending=True).reset_index(drop=True)\ntest_imageids = pd.Series(os.listdir(BASE_PATH + '/Test')).sort_values(ascending=True).reset_index(drop=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"cover_images_path = pd.Series(BASE_PATH + '/Cover/' + train_imageids ).sort_values(ascending=True)\nJMIPOD_images_path = pd.Series(BASE_PATH + '/JMiPOD/'+train_imageids).sort_values(ascending=True)\nJUNIWARD_images_path = pd.Series(BASE_PATH + '/JUNIWARD/'+train_imageids).sort_values(ascending=True)\nUERD_images_path = pd.Series(BASE_PATH + '/UERD/'+train_imageids).sort_values(ascending=True)\ntest_images_path = pd.Series(BASE_PATH + '/Test/'+test_imageids).sort_values(ascending=True)\nss = pd.read_csv(f'{BASE_PATH}/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"final=[]\ndef create_labels(cover,jmipod,juniward,uerd,image_id):\n    image = sk.imread(cover)\n    jmipodimg = sk.imread(jmipod)\n    juniward = sk.imread(juniward)\n    uerd = sk.imread(uerd)\n    \n    vec1 = np.reshape(image,(512*512*3))\n    vec2 = np.reshape(jmipodimg,(512*512*3))\n    vec3 = np.reshape(juniward,(512*512*3))\n    vec4 = np.reshape(uerd,(512*512*3))\n    \n    cos1 = spatial.distance.cosine(vec1,vec2)\n    cos2 = spatial.distance.cosine(vec1,vec3)\n    cos3 = spatial.distance.cosine(vec1,vec4)\n    \n    final.append({'image_id':image_id,'jmipod':cos1,'juniward':cos2,'uerd':cos3})","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I will be taking only the first 30k images , although you can all of the images if you want"},{"metadata":{"trusted":true},"cell_type":"code","source":"for k in tqdm(range(30000)):\n    create_labels(cover_images_path[k],JMIPOD_images_path[k],JUNIWARD_images_path[k],UERD_images_path[k],train_imageids[k])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_temp = pd.DataFrame(final)\ntrain_temp.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Adding softmax to our dissimilarity to get Probabilities"},{"metadata":{"trusted":true},"cell_type":"code","source":"def sigmoid(X):\n   return 1/(1+np.exp(-X))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_temp['jmipod'] = train_temp['jmipod'].apply(lambda x:sigmoid(x))\ntrain_temp['juniward'] = train_temp['juniward'].apply(lambda x:sigmoid(x))\ntrain_temp['uerd'] = train_temp['uerd'].apply(lambda x:sigmoid(x))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_temp.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"IMG_SIZE = 300\ndef load_training_data():\n  train_data = []\n  data_paths = [cover_images_path,JUNIWARD_images_path,JMIPOD_images_path,UERD_images_path]\n  labels = [np.zeros(train_temp.shape[0]),train_temp['juniward'],train_temp['jmipod'],train_temp['uerd']]\n  for i,image_path in enumerate(data_paths):\n    for j,img in enumerate(image_path[:10000]):\n        label = labels[i][j]\n        img = Image.open(img)\n        img = img.convert('L')\n        img = img.resize((IMG_SIZE, IMG_SIZE), Image.ANTIALIAS)\n        train_data.append([np.array(img), label])\n        \n  shuffle(train_data)\n  return train_data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def load_test_data():\n    test_data = []\n    for img in test_images_path:\n        img = Image.open(img)\n        img = img.convert('L')\n        img = img.resize((IMG_SIZE, IMG_SIZE), Image.ANTIALIAS)\n        test_data.append([np.array(img)])\n            \n    return test_data","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* I have used only top 44,000 images because if I run on more than that it exceeds memory allocation\n\n* Now that we have a data loader , we can now load the data and check if everything is working"},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntrain = load_training_data()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Sanity Check"},{"metadata":{"trusted":true},"cell_type":"code","source":"len(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.imshow(train[115][0], cmap = 'gist_gray')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"trainImages = np.array([i[0] for i in train]).reshape(-1, IMG_SIZE, IMG_SIZE, 1)\ntrainLabels = np.array([i[1] for i in train])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Model\n\nI am using 5 Conv2D layers with max pooling and batch norm for my baseline"},{"metadata":{"trusted":true},"cell_type":"code","source":"#PRELIMINARIES\nimport keras\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten\nfrom keras.layers import Conv2D, MaxPooling2D\nfrom keras.layers. normalization import BatchNormalization","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model = Sequential()\nmodel.add(Conv2D(32, kernel_size = (3, 3), activation='relu', input_shape=(IMG_SIZE, IMG_SIZE, 1)))\nmodel.add(MaxPooling2D(pool_size=(2,2)))\nmodel.add(BatchNormalization())\nmodel.add(Conv2D(64, kernel_size=(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2,2)))\nmodel.add(BatchNormalization())\nmodel.add(Conv2D(64, kernel_size=(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2,2)))\nmodel.add(BatchNormalization())\nmodel.add(Conv2D(96, kernel_size=(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2,2)))\nmodel.add(BatchNormalization())\nmodel.add(Conv2D(32, kernel_size=(3,3), activation='relu'))\nmodel.add(MaxPooling2D(pool_size=(2,2)))\nmodel.add(BatchNormalization())\nmodel.add(Dropout(0.2))\nmodel.add(Flatten())\nmodel.add(Dense(128, activation='relu'))\nmodel.add(Dense(1))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.compile(loss='mean_squared_error', optimizer='adam',metrics = ['mean_squared_error'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(model.summary())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"model.fit(trainImages, trainLabels, batch_size = 100, epochs = 3, verbose = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\ntest = load_test_data()\ntestImages = np.array([i[0] for i in test]).reshape(-1, IMG_SIZE, IMG_SIZE, 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"%%time\npredict = model.predict(testImages,batch_size=100)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ss['Label'] = predict","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ss.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_temp.to_csv('train.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}