{"cells":[{"metadata":{},"cell_type":"markdown","source":"<a class=\"anchor\" id=\"0\"></a>\n# **ALASKA2 : Image Steganalysis - All you need to know**\n\n\n\n## Introduction\n\n\nKaggle has recently launched a competition [ALASKA2 Image Steganalysis](https://www.kaggle.com/c/alaska2-image-steganalysis). \n\nBut wait, what is Image Steganalysis. So, the objective of this notebook is to describe [Image Steganalysis](https://en.wikipedia.org/wiki/Steganalysis) in detail so that Kaggle users can understand it.\n\nSo, let's get started.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"**I hope you find this notebook useful and your <font color=\"red\"><b>UPVOTES</b></font> keep me motivated**\n\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"<a class=\"anchor\" id=\"0.1\"></a>\n# **Table of Contents**\n\n\n1.\t[Introduction to Steganalysis](#1)\n2.\t[The difference between Steganalysis and Steganography](#2)\n3.\t[Steganographic methods](#3)\n    - 3.1\t[Spatial domain](#3.1)\n    - 3.2\t[Transform domain](#3.2)\n    - 3.3\t[Spread Spectrum Technique](#3.3)\n    - 3.4\t[Statistical Method](#3.4)\n    - 3.5\t[Distortion Techniques](#3.5)\n4.\t[Characteristics of a Strong Steganography method](#4)\n5.\t[Performance measure](#5)\n6.\t[Steganalysis approaches](#6)\n    - 6.1\t[Specific or Target Steganalysis](#6.1)\n    - 6.2\t[Blind or Universal Steganalysis](#6.2)\n7.\t[Steganalysis Methods and Techniques](#7)\n    - 7.1\t[Statistical Steganalysis](#7.1)\n    - 7.2\t[Steganalysis Machine Learning Techniques](#7.2)\n    - 7.3\t[Steganalysis Deep Learning Models](#7.3)\n8.\t[Steganalysis Tools](#8)\n9.\t[Study of Image Steganalysis Techniques](#9)\n10.\t[Applications of Steganalysis](#10)\n11.\t[Credits](#11)\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **1. Introduction to Steganalysis** <a class=\"anchor\" id=\"1\"></a>\n\n[Table of Contents](#0.1)\n\n\n[Steganalysis](https://en.wikipedia.org/wiki/Steganalysis) is the study of detecting hidden messages using [steganography](https://en.wikipedia.org/wiki/Steganography). This is analogous to [cryptanalysis](https://en.wikipedia.org/wiki/Cryptanalysis) applied to [cryptography](https://en.wikipedia.org/wiki/Cryptography). So, we need to know **steganagraphy**. The goal of steganalysis is to identify suspected packages, determine whether or not they have a payload encoded into them, and, if possible, recover that payload.\n\n\nUnlike cryptanalysis, in which intercepted data contains a message (though that message is encrypted), steganalysis generally starts with a pile of suspect data files, but little information about which of the files, if any, contain a payload. The steganalyst is usually something of a forensic statistician, and must start by reducing this set of data files (which is often quite large; in many cases, it may be the entire set of files on a computer) to the subset most likely to have been altered.\n\nBut, first let's visualize steganalysis. Visual representation of steganalysis is as follows :\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![Block diagram of Steganalysis](https://www.researchgate.net/profile/Bismita_Choudhury/publication/282889667/figure/fig1/AS:614435383152640@1523504219027/Block-diagram-of-Steganalysis.png)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **2. The difference between Steganalysis and Steganography** <a class=\"anchor\" id=\"2\"></a>\n\n\n[Table of Contents](#0.1)\n\n\nAs described earlier, **steganalysis** is the study of detecting hidden messages using **steganography**. So, we need to know **steganography**.\n\n\n[Steganography](https://en.wikipedia.org/wiki/Steganography) is the practice of concealing a file, message, image, or video within another file, message, image, or video. Generally, the hidden messages appear to be (or to be part of) something else: images, articles, shopping lists, or some other cover text. For example, the hidden message may be in invisible ink between the visible lines of a private letter. Some implementations of steganography that lack a shared secret are forms of security through obscurity. \n\nThe advantage of steganography over cryptography alone is that the intended secret message does not attract attention to itself as an object of scrutiny. Whereas [cryptography](https://en.wikipedia.org/wiki/Cryptography) is the practice of protecting the contents of a message alone, steganography is concerned both with concealing the fact that a secret message is being sent and its contents.\n\n\nThe difference between **steganalysis** and **steganagraphy** can be visualized as follows:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![Steganography and Steganalysis](https://www.researchgate.net/publication/3455264/figure/fig1/AS:669062505967617@1536528340179/Steganography-and-steganalysis.ppm)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"So, **steganography** can be used to transmit secret data without applying cryptography techniques throw network supports and trusted their security at the same time. So, **steganography** is defined as the art and science of hiding or embedding secret data through various multimedia containers or network supports such as videos, audio files, network packets, and last but not least digital images.\n\nThe summary of the evolution of the steganographic data carrier is presented in the figure down below.\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![Evolution of steganographic data carrier](https://miro.medium.com/max/1400/0*FlwlPZ1M0MyZDSqc.png)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **3. Steganographic methods** <a class=\"anchor\" id=\"3\"></a>\n\n\n[Table of Contents](#0.1)\n\n\nSteganographic methods can be classified mainly into five categories, although in some cases this classification is not possible:\n\n\n- Spatial domain\n- Transform domain\n- Spread spectrum\n- Statistical Methods\n- Distortion Methods\n\n\nNow, let's understand these methods in detail.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **3.1 Spatial Domain** <a class=\"anchor\" id=\"3.1\"></a>\n\n[Table of Contents](#0.1)\n\n\n**Spatial domain** steganographic techniques, also known as substitution techniques, are a group of relatively simple techniques that create a covert channel in the parts of the cover image in which changes are likely to be a bit scant when compared to the human’s eyes.\n\nOne of the ways to do so is to hide information in the Least Significant Bit (LSB) of the image data. This embedding method is basically based on the fact that the least significant bits in an image can be thought of as random noise, and consequently, they become not responsive to any changes on the image.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **3.2 Transform Domain** <a class=\"anchor\" id=\"3.2\"></a>\n\n[Table of Contents](#0.1)\n\n\nTransform domain embedding can be defined as a domain of embedding techniques for which a number of algorithms have been suggested. The process of embedding data in the frequency domain of a signal is much stronger than embedding principles that operate in the time domain.\n\nTransform domain techniques have an advantage over LSB techniques because they hide information in areas of the image that are less exposed to compression, cropping, and image processing.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **3.3 Spread Spectrum Technique** <a class=\"anchor\" id=\"3.3\"></a>\n\n[Table of Contents](#0.1)\n\n\nSpread spectrum transmission in radio communications transmits messages below the noise level for any given frequency. When employed with steganography, spread spectrum either deals with the cover image as noise or tries to add pseudo-noise to the cover image.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **3.4 Statistical Method** <a class=\"anchor\" id=\"3.4\"></a>\n\n[Table of Contents](#0.1)\n\nAlso known as model-based techniques, these techniques tend to modulate or modify the statistical properties of an image in addition to preserving them in the embedding process. This modification is typically small, and it is thereby able to take advantage of the human weakness in detecting luminance variation.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **3.5 Distortion Techniques** <a class=\"anchor\" id=\"3.5\"></a>\n\n[Table of Contents](#0.1)\n\n\nDistortion techniques require knowledge of the original cover image during the decoding process where the decoder functions to check for differences between the original cover image and the distorted cover image in order to restore the secret message. The encoder, on the other hand, adds a sequence of changes to the cover image. So, information is described as being stored by signal distortion.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **4. Characteristics of a Strong Steganography method** <a class=\"anchor\" id=\"4\"></a>\n\n\n[Table of Contents](#0.1)\n\n\nThough steganography’s most obvious goal is to hide data, there are several other related goals used to judge a method’s steganographic strength. These include:\n\n\n- **Capacity** (how much data can be hidden)\n- **Invisibility** (inability for humans to detect a distortion in the stego-object)\n- **Undetectability** (inability for a computer to use statistics or other computational methods to differentiate between covers and stego-objects)\n- **Robustness** (message’s ability to persist despite compression or other common modifications)\n- **Tamper resistance** (message’s ability to persist despite active measures to destroy it)\n- **Signal to noise ratio (SNR)** (how much data is encoded versus how much-unrelated data is encoded).\n\n\nThe three main components, which work in opposition to one another, are **capacity**, **undetectability** and **robustness**.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **5. Performance measure** <a class=\"anchor\" id=\"5\"></a>\n\n[Table of Contents](#0.1)\n\nAs a performance measure for image distortion due to embedding, the well-known **peak-signal-to-noise ratio (PSNR)**, which is categorized under difference distortion metrics, can be applied to stego images. It is defined as:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![PSNR](https://miro.medium.com/max/160/0*0VSzwVHAXkYN7SkI.png)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"In the previous equation, **R is the maximum fluctuation in the input image data type**. For example, if the input image has a double-precision floating-point data type, then R is 1. If it has an 8-bit unsigned integer data type, R is 255, etc. **MSE denotes the mean square error**, which is given as","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![MSE](https://miro.medium.com/max/201/0*47pXS4-c7xer-mRf.png)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"M and N are the number of rows and columns in the input image.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The MSE represents the cumulative squared error between the stego and the cover image, whereas PSNR represents a measure of the peak error. The lower the value of MSE, the lower the error.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The **Structural Similarity (SSIM) Index quality assessment index** is based on the computation of three terms, namely the luminance term, the contrast term, and the structural term. The overall index is a multiplicative combination of the three terms.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![SSIM](https://miro.medium.com/max/261/0*eREOtTZLsdQKGngR.png)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **6. Steganalysis approaches** <a class=\"anchor\" id=\"6\"></a>\n\n[Table of Contents](#0.1)\n\n\nThe steganalysis algorithm may or may not depend on the **steganographic algorithm (SA)**. Based on this, steganalysis is classified as follows:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **6.1 Specific or Target steganalysis** <a class=\"anchor\" id=\"6.1\"></a>\n\n[Table of Contents](#0.1)\n\nThe SA is known and the designing of detector (steganalysis algorithm) is based on SA. The steganalysis algorithm is dependent on the SA. This type of steganalysis is based on analyzing the statistical properties of an image that change after embedding. The advantage of using specific steganalysis is the results are very accurate. The disadvantage of using this method is it is very limited to particular embedding algorithm as well as the image format.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **6.2 Blind or Universal steganalysis** <a class=\"anchor\" id=\"6.2\"></a>\n\n[Table of Contents](#0.1)\n\nIn universal steganalysis, the SA is not known by everyone. Hence, anyone can design a detector to detect the presence of the secret message that will not depend on SA. Comparing with specific steganalysis, universal is common and less efficient. Still universal steganalysis is widely used than specific one because it is independent of the SA. **Nowadays, scientific research focuses on universal steganalysis**. It is visually represented as follows:","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"![Image Steganalysis](https://miro.medium.com/max/2800/0*nCbBS1_pq9N5nDh4.png)","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# **7. Steganalysis Methods and Techniques** <a class=\"anchor\" id=\"7\"></a>\n\n[Table of Contents](#0.1)\n\n\nBased on the way of detecting the presence of a hidden message, steganalysis methods are divided as follows:","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **7.1 Statistical Steganalysis** <a class=\"anchor\" id=\"7.1\"></a>\n\n[Table of Contents](#0.1)\n\nIn order to detect the existence of the hidden message, statistical analysis is done with the pixels. It is further classified as **spatial domain** steganalysis and **transform domain** steganalysis.\n\nIn the spatial domain, the pair of pixels is considered and the difference between them is calculated. The pair may be any 2 neighboring pixels. They may be selected within a block otherwise across the two blocks. Finally, the histogram is plotted that shows the existence of the hidden message.\n\nIn the transform domain, frequency counts of coefficients are calculated and then histogram analysis is performed. With the help of this, the cover and stego images can be differentiated. However, this method is not providing information about the embedding algorithms. To overcome this problem, we may choose feature based steganalysis.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **7.2 Steganalysis Machine Learning Techniques** <a class=\"anchor\" id=\"7.2\"></a>\n\n[Table of Contents](#0.1)\n\n\nIn this method, the features of the image will be extracted for selecting and retaining relevant information. These extracted features are used to detect a hidden message in an image. They can also be used to train classifiers.\n\n\nIt includes the following 2 phases:\n\n\n- **a. Feature Extraction**: It is a process of creating a set of distinct statistical attributes of an image. These attributes are known as a feature. Feature Extraction is nothing but a dimensionality reduction. The extracted features must be sensitive to the embedding artifacts. Image quality metrics, wavelet decompositions, moment of image statistic histograms, Markov empirical transition matrix, moment of image statistic from spatial and frequency domain, co-occurrence matrix are some of the feature extraction methods.\n\n\n- **b. Classification**: It is a way of categorizing the images into classes depending on their feature values. Supervised learning is one of the primary classifications in steganalysis. Supervised learning allows learning under some supervision. In this learning, a set of training inputs that includes input features is given as input to train the classifier. After the training, class label is predicted based on the features that are given.\n\n\nsteganalysis use the following classifiers:\n\n\n- **Multivariate regression**: It consists of the regression coefficient. In the training phase, regression coefficients are predicted using minimum mean square error.\n\n- **FLD**: It is a linear combination of features which maximizes the separations. In the classification method, multidimensional features are projected into a linear space.\n\n- **SVM**: This classification method learns from the given sample. It is trained to recognize and assign class labels based on a given set of features.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"## **7.3 Steganalysis Deep Learning Models** <a class=\"anchor\" id=\"7.3\"></a>\n\n[Table of Contents](#0.1)\n\nIn recent years, Deep Learning models such as the Convolutional Neural Networks (CNN or ConvNet) have been introduced and applied within image steganalysis context. Deep learning combines the two phases in machine learning techniques(Feature extraction and classification) as a single compact model.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **8. Steganalysis Tools** <a class=\"anchor\" id=\"8\"></a>\n\n[Table of Contents](#0.1)\n\nVarious steganalysis tools are available to detect the presence of hidden information with the stego image. Some of the steganalysis tools are mentioned below:\n\n1. [StegDetect](https://github.com/abeluck/stegdetect)\n2. [Stirmark](https://www.petitcolas.net/watermarking/stirmark/)\n3. [StegBreak](https://github.com/s-fiebig/stegdetect-stegbreak)\n4. [StegSecret](https://www.aldeid.com/wiki/StegSecret)\n5. [JPSeek](https://github.com/h3xx/jphs)\n6. [2Mosaic](http://www.petitcolas.net/watermarking/2mosaic/)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **9. Study of Image Steganalysis Techniques** <a class=\"anchor\" id=\"9\"></a>\n\n[Table of Contents](#0.1)\n\n\nAlthough steganography may provide a safe tool for communication to government and business, it will suffer serious consequences if it is used by the terrorists or criminals. In contrast, steganalysis was proposed to determine whether there are secret messages embedded in image or video. \n\n\nSteganalysis involves two major types of analysis: **Visual analysis** and **statistical analysis**. Visual analysis deals with detection of secret message with naked eye or with the help of computer in which bit planes are analyzed separately for any unusual change in the appearance for the presence of secret message. Statistical analysis deals with detection of any change in statistical properties of stego object caused by steganographic algorithm. Steganalysis can be divided into two major types: Universal steganalysis and specific type of steganalysis techniques. Universal steganalysis techniques can detect secret message in stego objects embedded by a range of steganographic algorithms and specific stegalalysis techniques, which are more sophisticated techniques and work corresponding to a particular steganographic algorithm only.\n\nUniversal steganalysis is also referred as blind steganalysis because these are independent of any specific embedding technique and are used to alleviate the deficiency of targeted analyzers by removing their dependency on the behavior of individual embedding techniques. To achieve this, a set of distinguishing statistics that are sensitive to a wide variety of embedding operations are determined and collected. These statistics, computed from both cover and stego images are employed to train a classifier, which is subsequently used to distinguish between cover and stego images.\n\nBlind steganalysis has two important components; these are **feature extraction** and **feature classification**. In feature extraction, a set of distinguishing statistics are obtained from a data set of images. There is no well defined approach for obtaining these statistics, but often they are proposed by observing general image features that show inputs are the images and the outputs are the class labels. Feature extraction is often used in blind steganalysis which aims to build a universal classifier to detect all steganography without knowing the actual methods used. Classifiers are used to wrap the extracted feature. The classifier is designed using the training set, some classified instances are applied, and then the classifier is used to classify unknown test set. A variety of classification methods can be used and different classification algorithms can be designed for better performance of the steganalyzer. For steganalysis, the spatial domain & transform domain features extracted from the stego image could be applied to the neural network classifier in order to classify the original and the stego image. Suitable learning machine may be applied as a classifier for the desired steganalysis process. The general computational intelligence based approach for steganalysis can be illustrated as follows -","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"![Image Steganalysis Techniques](https://www.spiedigitallibrary.org/ContentImages/Journals/JEIME5/20/1/013016/WebImages/013016_1_1.jpg)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **10. Applications of Steganalysis** <a class=\"anchor\" id=\"10\"></a>\n\n[Table of Contents](#0.1)\n\n\n- **Medical safety**: Current image formats such as DICOM separate image data from the text (such as patients name, date and physician), with the result that the link between image and patient occasionally gets mangled by protocol converters. Thus embedding the patients name in the image could be a useful safety measure.\n\n- **Intellectual property offenses**: Intellectual property, defined as the formulas, prototypes, copyrights and customer lists maintained by a company, can be far more valuable than the actual items they sell.\n\n- **Watermarking**: Special inks to write hidden messages on bank notes and also the entertainment industry using digital watermarking and fingerprinting of audio and video for copyright protection.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# **11. Credits** <a class=\"anchor\" id=\"11\"></a>\n\n[Table of Contents](#0.1)\n\n\n- https://en.wikipedia.org/wiki/Steganalysis\n\n- https://medium.com/@rabi3elbeji/friendly-introduction-to-steganalysis-716d9147c17c\n\n- https://medium.com/@rabi3elbeji/friendly-introduction-to-steganography-4cf032240904\n\n- https://shodhganga.inflibnet.ac.in/bitstream/10603/8912/13/11_chapter%202.pdf\n\n","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"[Go to Top](#0)","execution_count":null}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}