{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":5127,"databundleVersionId":868727,"sourceType":"competition"},{"sourceId":10224995,"sourceType":"datasetVersion","datasetId":6321340}],"dockerImageVersionId":30787,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"text-align: center;\">\n  <h1><b>Neural Style Transfer: Leveraging VGG19 to Create Painting-Like Images</b></h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### By: Philip Gort, Ralph ten Broek and Charlotte Groot - Group 3","metadata":{}},{"cell_type":"markdown","source":"# Table of Contents \n* [1. Introduction](#h1)\n    * [1.1. Outline of the Notebook](#11)\n    * [1.2. Imports](#12)\n* [2. Methods](#2)\n    * [2.1. Data Import](#21)\n    * [2.2. Preprocessing Data](#22)\n    * [2.3. Neural Style Transfer](#23)\n        * [2.3.1. Insert Image](#231)\n        * [2.3.2. VGG19](#232)\n        * [2.3.3. New Model](#233)\n* [3. Results](#3)\n* [4. Conclusion/ Discussion](#4)\n    * [4.1. Generalizability](#41)\n    * [4.2. Opportunities for Improvement](#42)\n    * [4.3. Important Takeaways for Practitioners](#43)\n    * [4.4. Important Takeaways for Researchers](#44)\n    * [4.5. Negation Words](#45)\n* [5. References](#5)\n* [6. Division of Labour](#6)","metadata":{}},{"cell_type":"markdown","source":"# 1. Introduction <a class=\"anchor\"  id=\"h1\"></a>\n\nFor this project, we aimed to combine art and technology by applying neural style transfer to transform images into paintings. Style transfer is a computer vision technique that generates images with a new style by combining the content of one image with the style of another (Xu, 2024).\n\nOur research question is: How do the results differ between using a single style image compared to multiple style images for generating an artistic image? To answer this question, we used a pretrained VGG19 network, a very useful tool for extracting image features. We focused on paintings of the 4 most common styles in our dataset: Romanticism, Impressionism, Realism, and Expressionism.\n\nOur work is relevant for both artists and researchers. For artists, it provides a tool for creating unique artworks by blending  styles, enabling creative exploration and innovation in digital art. Subsequently, the generated images have potential applications in advertisement, media, and entertainment. For researchers in deep learning, our work contributes to understanding style transfer techniques, particularly the impact of using single versus multiple style images, which can inform future research.\n\nOur goal is to deepen our understanding of neural style transfer and explore trade-offs between computational efficiency and visual quality. We used Xu's paper as inspiration, who also applied neural style transfer (2024). Additionally, we used several online tutorials (fchollet, 2020; Telega, 2024).","metadata":{}},{"cell_type":"markdown","source":"## 1.1 Outline of the Notebook <a class=\"anchor\"  id=\"11\"></a>\n\nThis notebook is structured as follows: we will load and preprocess the dataset. Next, we use the VGG19 network to extract features, after which we will apply style transfer using either 1 or 5 style images. We will visualise our results, report the content and style loss, and evaluate performance using the Structural Similarity Index (SSIM). Finally, we will compare the results, where we will discuss generalizability, suggest opportunities for improvement and highlight important takeaways.","metadata":{}},{"cell_type":"markdown","source":"## 1.2 Imports <a class=\"anchor\"  id=\"12\"></a>\n\nBefore we start importing the dataset, we load the necessary packages and libraries:","metadata":{}},{"cell_type":"code","source":"# Standard Python libraries\nimport os\nimport zipfile\nimport re\nimport itertools\n\n# Numerical and data manipulation libraries\nimport numpy as np\nimport pandas as pd\n\n# TensorFlow and Keras libraries for deep learning\nimport tensorflow as tf\nfrom tensorflow.keras.applications import VGG19\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.preprocessing.image import load_img, img_to_array\n\n# Image processing and computer vision libraries\nimport cv2\nfrom PIL import Image\n\n# Plotting and visualization libraries\nimport matplotlib\nimport matplotlib.pyplot as plt\nfrom mpl_toolkits.axes_grid1 import make_axes_locatable\nfrom matplotlib.colors import ListedColormap, LinearSegmentedColormap\n\n# Machine learning tools for data splitting\nfrom sklearn.model_selection import train_test_split\n\n# Image feature extraction (e.g. HOG)\nfrom skimage.feature import hog\n\n# Image similarity metric (Structural Similarity Index)\nfrom skimage.metrics import structural_similarity as compare_ssim\n\n# Data visualization and exploration\nimport seaborn as sns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:42:08.156638Z","iopub.execute_input":"2024-12-22T16:42:08.156896Z","iopub.status.idle":"2024-12-22T16:42:20.664036Z","shell.execute_reply.started":"2024-12-22T16:42:08.156869Z","shell.execute_reply":"2024-12-22T16:42:20.663325Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 2. Methods <a class=\"anchor\"  id=\"h2\"></a>\n\nThe dataset we used was compiled by Kaggle and based on the WikiArt dataset. This dataset includes approximately 100,000 paintings from 2,300 artists, covering a wide range of time periods and artistic styles. We will extract style features from the paintings in this dataset, enabling our style transfer model to apply these styles when transforming images.","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Data Import <a class=\"anchor\"  id=\"21\"></a>","metadata":{}},{"cell_type":"code","source":"# Define the paths of the files we want\nfile_paths = [\n    '/kaggle/input/painter-by-numbers/all_data_info.csv',\n    '/kaggle/input/painter-by-numbers/train_1.zip'\n]\n\n# Check if the files exist and print their paths\nfor path in file_paths:\n    if os.path.exists(path):\n        print(f\"Found file: {path}\")\n    else:\n        print(f\"File not found: {path}\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:42:20.666062Z","iopub.execute_input":"2024-12-22T16:42:20.666651Z","iopub.status.idle":"2024-12-22T16:42:20.673189Z","shell.execute_reply.started":"2024-12-22T16:42:20.666606Z","shell.execute_reply":"2024-12-22T16:42:20.672350Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The metadata also has to be imported. This contains the specifications of the images, like: \"year\", \"artist\", \"style\", etc.","metadata":{}},{"cell_type":"code","source":"artistdata = pd.read_csv('/kaggle/input/painter-by-numbers/all_data_info.csv')\nartistdata.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:42:20.675160Z","iopub.execute_input":"2024-12-22T16:42:20.675611Z","iopub.status.idle":"2024-12-22T16:42:21.129964Z","shell.execute_reply.started":"2024-12-22T16:42:20.675553Z","shell.execute_reply":"2024-12-22T16:42:21.129107Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.2 Preprocessing Data <a class=\"anchor\"  id=\"22\"></a>\n\nTo unpack the data, we use both files: one containing the paths to the images and another containing the metadata associated with those images. Both files are essential for this project.","metadata":{}},{"cell_type":"code","source":"# Open and extract ZipFile\nnp.random.seed(42) \nimport zipfile\nwith zipfile.ZipFile('/kaggle/input/painter-by-numbers/train_1.zip', 'r') as zip_ref:\n    zip_ref.extractall('/kaggle/working/train_data')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:42:21.130998Z","iopub.execute_input":"2024-12-22T16:42:21.131350Z","iopub.status.idle":"2024-12-22T16:43:08.970856Z","shell.execute_reply.started":"2024-12-22T16:42:21.131311Z","shell.execute_reply":"2024-12-22T16:43:08.970096Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Directory where images were extracted\ntrain_data_dir = '/kaggle/working/train_data/train_1'\n\n# List all files in the training directory\nimage_files = [os.path.join(train_data_dir, file) for file in os.listdir(train_data_dir)]\n\n# Load and display the image inline\nimage_path = image_files[2]\nimg = Image.open(image_path)\n\n# Display using matplotlib\nplt.imshow(img)\nplt.axis('off')  # Hide axes for cleaner display\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:08.971922Z","iopub.execute_input":"2024-12-22T16:43:08.972168Z","iopub.status.idle":"2024-12-22T16:43:09.359629Z","shell.execute_reply.started":"2024-12-22T16:43:08.972145Z","shell.execute_reply":"2024-12-22T16:43:09.358407Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"From exploratory analysis, we found that \"Impressionism\", \"Realism\", \"Romanticism\" and \"Expressionism\" are the top 4 most present styles in the dataset. We will only select paintings from these styles, for two reasons: to avoid an overly large dataset (which could lead to longer computation times and higher memory usage) and to ensure that each style is sufficiently represented (allowing for meaningful feature extraction). Eventually, 4,005 paintings were selected, each labeled in the meta dataset with its corresponding style.","metadata":{}},{"cell_type":"code","source":"top4styles = ['Impressionism', 'Realism', 'Romanticism', 'Expressionism']\nartistdata = artistdata[artistdata['style'].isin(top4styles)]\nprint(artistdata)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:09.360805Z","iopub.execute_input":"2024-12-22T16:43:09.361055Z","iopub.status.idle":"2024-12-22T16:43:10.647652Z","shell.execute_reply.started":"2024-12-22T16:43:09.361029Z","shell.execute_reply":"2024-12-22T16:43:10.646697Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here, the metadata will be connected to the corresponding images.","metadata":{}},{"cell_type":"code","source":"train_metadata = artistdata[artistdata['in_train'] == True]\ntrain_filenames = train_metadata['new_filename'].tolist()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:10.650299Z","iopub.execute_input":"2024-12-22T16:43:10.650597Z","iopub.status.idle":"2024-12-22T16:43:10.658135Z","shell.execute_reply.started":"2024-12-22T16:43:10.650542Z","shell.execute_reply":"2024-12-22T16:43:10.657419Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"filtered_image_files = [\n    file for file in image_files if os.path.basename(file) in train_filenames\n]\n\n# Final filtering of artist data to match the filtered image files\nfinal_filenames = [os.path.basename(file) for file in filtered_image_files]\nfiltered_artist_data = artistdata[artistdata['new_filename'].isin(final_filenames)]\nprint(f\"Final filtered artist data shape: {filtered_artist_data.shape}\")\n\nprint(f\"Final count of filtered image files: {len(filtered_image_files)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:10.659134Z","iopub.execute_input":"2024-12-22T16:43:10.659415Z","iopub.status.idle":"2024-12-22T16:43:14.880368Z","shell.execute_reply.started":"2024-12-22T16:43:10.659389Z","shell.execute_reply":"2024-12-22T16:43:14.879470Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Number of unique styles\nunique_styles = filtered_artist_data['style'].nunique()\nprint(f\"Number of unique styles: {unique_styles}\")\n\n# Number of paintings per style\npaintings_per_style = filtered_artist_data['style'].value_counts()\nprint(\"\\nNumber of paintings per style:\")\nprint(paintings_per_style)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:14.881536Z","iopub.execute_input":"2024-12-22T16:43:14.881832Z","iopub.status.idle":"2024-12-22T16:43:14.892362Z","shell.execute_reply.started":"2024-12-22T16:43:14.881806Z","shell.execute_reply":"2024-12-22T16:43:14.891430Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Initially, our goal was to let the user choose between the 4 styles for their image transformation. However, due to Kaggle's memory limitations, we were constrained in our data options and model complexity, which forced us to alter the direction of our project.\n\nFor our initial purpose, we already splitted the dataset into training, validation, and test sets using a 80-10-10 split per style. This was done by the grouped variable. \n\nAlthough this step is not necessary for the purpose of our adjusted objective, we still decided to keep our code, since it could be useful for future research on our initial goal.","metadata":{}},{"cell_type":"code","source":"train_art = pd.DataFrame()\ntest_art = pd.DataFrame()\nval_art = pd.DataFrame()\n\n# Group by style\ngrouped = filtered_artist_data.groupby('style')\n\n# Check the number of styles and the distribution of paintings per style\nprint(f\"Number of styles: {len(grouped)}\")\nprint(f\"Paintings per style:\\n{filtered_artist_data['style'].value_counts()}\")\n\n# Loop through each style group\nfor style, group in grouped:\n\n    shuffled_group = group.sample(frac=1, random_state=2024).reset_index(drop=True)\n\n    # Split into train, validation, and test\n    train, temp = train_test_split(shuffled_group, test_size=0.2, random_state=2024)\n    val, test = train_test_split(temp, test_size=0.5, random_state=2024)\n\n    # Append to the corresponding datasets\n    train_art = pd.concat([train_art, train], ignore_index=True)\n    val_art = pd.concat([val_art, val], ignore_index=True)\n    test_art = pd.concat([test_art, test], ignore_index=True)\n\n# Debug output for the sizes of the datasets\nprint(f\"Training set size: {train_art.shape[0]}\")\nprint(f\"Validation set size: {val_art.shape[0]}\")\nprint(f\"Test set size: {test_art.shape[0]}\")\n\n# Check if the final size of the training, validation, and test sets match the expected proportions\ntotal_images = len(filtered_artist_data)\nprint(f\"Total number of filtered images: {total_images}\")\nprint(f\"Expected Training set size: {0.8 * total_images}\")\nprint(f\"Expected Validation set size: {0.1 * total_images}\")\nprint(f\"Expected Test set size: {0.1 * total_images}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:14.893376Z","iopub.execute_input":"2024-12-22T16:43:14.893689Z","iopub.status.idle":"2024-12-22T16:43:14.926222Z","shell.execute_reply.started":"2024-12-22T16:43:14.893646Z","shell.execute_reply":"2024-12-22T16:43:14.925328Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Link it so that we also have the split among the images\ntrain_img_paths = [img for img in filtered_image_files if os.path.basename(img) in train_art['new_filename'].tolist()]\nval_img_paths = [img for img in filtered_image_files if os.path.basename(img) in val_art['new_filename'].tolist()]\ntest_img_paths = [img for img in filtered_image_files if os.path.basename(img) in test_art['new_filename'].tolist()]\nprint(f\"Number of train file paths: {len(train_img_paths)}\")\nprint(f\"Number of train file paths: {len(val_img_paths)}\")\nprint(f\"Number of train file paths: {len(test_img_paths)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:14.927746Z","iopub.execute_input":"2024-12-22T16:43:14.928113Z","iopub.status.idle":"2024-12-22T16:43:15.298078Z","shell.execute_reply.started":"2024-12-22T16:43:14.928075Z","shell.execute_reply":"2024-12-22T16:43:15.297149Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For a more generalized model, it is important that all images have the same size. Furthermore, normalization of the values is important. It ensures that the inputs to different layers of a neural network have similar distributions, which can significantly improve the model's performance and stability. \n\nWe also introduced data augmentation in the forms of random crop, size, contrast, brightness, horizontal flip to ensure better generalization. While center cropping and augmentation are not necessary for the adjusted objective, we again decided to keep these steps in our code for future research.\n\n*This code section was generated with the help of ChatGPT.*","metadata":{}},{"cell_type":"code","source":"# Preprocess image by applying center crop, resizing, and normalizing pixel values\ndef preprocess_image_with_center_crop(image_path, crop_size, target_size):\n    try:\n        img = Image.open(image_path).convert(\"RGB\")  # Convert to RGB\n        width, height = img.size\n        \n        # Calculate cropping box for center crop\n        crop_width, crop_height = crop_size\n        left = (width - crop_width) // 2\n        top = (height - crop_height) // 2\n        right = left + crop_width\n        bottom = top + crop_height\n        \n        # Perform center crop\n        img_cropped = img.crop((left, top, right, bottom))\n        \n        # Resize the cropped image to the target size\n        img_resized = img_cropped.resize(target_size)\n        \n        # Normalize pixel values to [0, 1]\n        img_normalized = np.array(img_resized) / 255.0\n        \n        return img_normalized\n    except Exception as e:\n        print(f\"Error processing image {image_path}: {e}\")\n        return None\n\ncenter_crop_size = (299, 299)  \nimage_size = (299, 299)   \n\n# Define augment_image function\ndef augment_image(image):\n    image = tf.image.random_flip_left_right(image)  # Random horizontal flip\n    image = tf.image.random_brightness(image, max_delta=0.2)  # Random brightness\n    image = tf.image.random_contrast(image, lower=0.8, upper=1.2)  # Random contrast\n    image = tf.image.random_saturation(image, lower=0.8, upper=1.2)  # Random saturation\n    image = tf.image.random_crop(image, size=(250, 250, 3))  # Random crop\n    return tf.image.resize(image, (299, 299))  # Resize back to target size\n\n# Update preprocessing for datasets and apply data augmentation for training set\ndef preprocess_dataset_with_center_crop(file_list, augment=False):\n    preprocessed_images = []\n    for file in file_list:\n        # Preprocess image: center crop and resize\n        processed_img = preprocess_image_with_center_crop(file, crop_size=center_crop_size, target_size=image_size)\n        # Convert the image to a TensorFlow tensor\n        processed_img = tf.convert_to_tensor(processed_img, dtype=tf.float32)\n        # Apply augmentation if specified (only for training set)\n        if augment:\n            processed_img = augment_image(processed_img)\n        preprocessed_images.append(processed_img)\n    return np.array(preprocessed_images)\n\n# Apply preprocessing to datasets with augmentation for training data\nX_train = preprocess_dataset_with_center_crop(train_img_paths, augment=True)  # Augment for training\nX_val = preprocess_dataset_with_center_crop(val_img_paths, augment=False)    # No augmentation for validation\nX_test = preprocess_dataset_with_center_crop(test_img_paths, augment=False)  # No augmentation for test\n\n# Verify shapes\nprint(f\"Training set shape: {X_train.shape}\")\nprint(f\"Validation set shape: {X_val.shape}\")\nprint(f\"Test set shape: {X_test.shape}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:43:15.299474Z","iopub.execute_input":"2024-12-22T16:43:15.299845Z","iopub.status.idle":"2024-12-22T16:44:34.388305Z","shell.execute_reply.started":"2024-12-22T16:43:15.299805Z","shell.execute_reply":"2024-12-22T16:44:34.387356Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"processed_image = X_val[10]\n\n# Rescale the image back to the [0, 255] range and convert to uint8\nprocessed_image_rescaled = (processed_image * 255).astype(np.uint8)\n\n# Convert the rescaled image back to a PIL image\nprocessed_img_pil = Image.fromarray(processed_image_rescaled)\n\n# Display the processed image\nplt.imshow(processed_img_pil)\nplt.axis('off')  # Hide axes for cleaner display\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:34.389269Z","iopub.execute_input":"2024-12-22T16:44:34.389511Z","iopub.status.idle":"2024-12-22T16:44:34.522260Z","shell.execute_reply.started":"2024-12-22T16:44:34.389488Z","shell.execute_reply":"2024-12-22T16:44:34.521436Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2.3 Neural Style Transfer <a class=\"anchor\"  id=\"23\"></a>\n\nNow that we have preprocessed our dataset, we are ready to build a model for performing neural style transfer.","metadata":{}},{"cell_type":"markdown","source":"### 2.3.1 Insert Image <a class=\"anchor\"  id=\"211\"></a>\n\nIn this step, we specify the image we want to transfer, called the \"Content Image\". For demonstration purposes, we picked a random image of a dog. However, the image paths can be replaced with any image. ","metadata":{}},{"cell_type":"code","source":"hond = \"/kaggle/input/hondpath/doggo.webp\"\n\n# Show content image\nimg = Image.open(hond)\nplt.imshow(img)\nplt.title(\"Content Image\")\nplt.axis('off')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:34.523222Z","iopub.execute_input":"2024-12-22T16:44:34.523552Z","iopub.status.idle":"2024-12-22T16:44:34.836902Z","shell.execute_reply.started":"2024-12-22T16:44:34.523516Z","shell.execute_reply":"2024-12-22T16:44:34.836037Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.3.2 VGG19 <a class=\"anchor\"  id=\"232\"></a>\n\nIn this project we used VGG19 to extract features. VGG19 is a pretrained image classification network that is highly suitable for neural style transfer due to its ability to define both content and style representations from images (Xu, 2024). \n\nFirst we will define the constants. The weights will balance the importance of style versus content in transforming the image. ","metadata":{}},{"cell_type":"code","source":"# VGG19 layer to use for content representation\nCONTENT_LAYER = 'block5_conv2'  \n# List of VGG19 layers used to extract style features (e.g. textures, patterns)\nSTYLE_LAYERS = ['block1_conv1', 'block2_conv1', 'block3_conv1', 'block4_conv1', 'block5_conv1']\n\n# Weight assigned to the style loss\nSTYLE_WEIGHT = 1\n# Weight assigned to the content loss\nCONTENT_WEIGHT = 10\n# Dynamic variable for color weight\nCOLOR_WEIGHT = tf.Variable(.5, dtype=tf.float32, trainable=False)  ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:34.838027Z","iopub.execute_input":"2024-12-22T16:44:34.838373Z","iopub.status.idle":"2024-12-22T16:44:34.847533Z","shell.execute_reply.started":"2024-12-22T16:44:34.838335Z","shell.execute_reply":"2024-12-22T16:44:34.846708Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Here, we define two functions: one to load and preprocess the images, and another to deprocess the model's output back to a displayable RGB image. \n\nThe numbers 103.939, 116.779, and 123.68 in the deprocess_image function are the mean pixel values for the RGB channels of the ImageNet dataset, on which a VGG19 is trained. These values normalize input images which helps with consistency with the data the model was trained on. At the deprocessing step, we add back these means to change the image back to its original colours. \n\n*This code section was build with the help of ChatGPT.*","metadata":{}},{"cell_type":"code","source":"# Function to load and preprocess images\ndef load_and_process_image(image_path, target_size=(299, 299)):\n    img = load_img(image_path, target_size=target_size) # Load and resize the image\n    img = img_to_array(img) # Convert image to numpy array\n    img = np.expand_dims(img, axis=0) # Add batch dimension (model expects inputs in batches)\n    img = tf.keras.applications.vgg19.preprocess_input(img) # Normalize image for VGG19 model\n    return img\n\n# Function to deprocess image\ndef deprocess_image(img):\n    img = img.reshape((299, 299, 3))\n    img[:, :, 0] += 103.939 # Add back the mean values subtracted during preprocessing.\n    img[:, :, 1] += 116.779\n    img[:, :, 2] += 123.68\n    img = img[:, :, ::-1]\n    img = np.clip(img, 0, 255).astype('uint8')\n    return img","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:34.848581Z","iopub.execute_input":"2024-12-22T16:44:34.848847Z","iopub.status.idle":"2024-12-22T16:44:34.860171Z","shell.execute_reply.started":"2024-12-22T16:44:34.848823Z","shell.execute_reply":"2024-12-22T16:44:34.859481Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Next, we specify the paths to the images we picked. The neural style transfer will use the style of either 1 or 5 of these images. Again, the image paths can be replaced with any images that are similair to each other. After browsing, we picked images of the style \"impressionism\". Impressionism is characterized by the short brush strokes and the use of natural lights and colours, as can be seen in the paintings below (Painting, 2022). These are the features we are interested in transfering.","metadata":{}},{"cell_type":"code","source":"# Load content and style images\ncontent_image_path = hond  \n\n# Loading in the style images we picked by browsing.\nstyle_image_paths = [\n    '/kaggle/working/train_data/train_1/1642.jpg',\n    '/kaggle/working/train_data/train_1/16150.jpg',\n    '/kaggle/working/train_data/train_1/102363.jpg',\n    '/kaggle/working/train_data/train_1/10297.jpg',\n    '/kaggle/working/train_data/train_1/13960.jpg'\n]\n\ncontent_image = load_and_process_image(content_image_path)\nstyle_image = load_and_process_image(style_image_paths[1])\nstyle_images = [load_and_process_image(path) for path in style_image_paths]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:34.861435Z","iopub.execute_input":"2024-12-22T16:44:34.861791Z","iopub.status.idle":"2024-12-22T16:44:34.924010Z","shell.execute_reply.started":"2024-12-22T16:44:34.861754Z","shell.execute_reply":"2024-12-22T16:44:34.923086Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for path in style_image_paths:\n    img = Image.open(path)\n    plt.imshow(img)\n    plt.title(\"Style Image\")\n    plt.axis('off')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:34.925038Z","iopub.execute_input":"2024-12-22T16:44:34.925279Z","iopub.status.idle":"2024-12-22T16:44:36.235081Z","shell.execute_reply.started":"2024-12-22T16:44:34.925256Z","shell.execute_reply":"2024-12-22T16:44:36.234315Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now, we build the VGG19 model to extract features from the images above. This convolutional network consists of 19 weight layers, of which 16 convolutional layers and 3 fully connected layers (Xu, 2024). We froze the entire network, meaning we did not fine-tune the VGG19 model.","metadata":{}},{"cell_type":"code","source":"# Load pre-trained VGG19 model without classification head\nvgg = VGG19(include_top=False, weights='imagenet')\n# Freeze the model so it’s used only for feature extraction\nvgg.trainable = False","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:36.236330Z","iopub.execute_input":"2024-12-22T16:44:36.237038Z","iopub.status.idle":"2024-12-22T16:44:37.189483Z","shell.execute_reply.started":"2024-12-22T16:44:36.236998Z","shell.execute_reply":"2024-12-22T16:44:37.188793Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.3.3 New Model <a class=\"anchor\"  id=\"233\"></a>","metadata":{}},{"cell_type":"markdown","source":"Here, we create a function that creates a new model based on the VGG19 network, which extracts outputs from the style and content layers. ","metadata":{}},{"cell_type":"code","source":"# Extract style and content layers\ndef get_model():\n    style_outputs = [vgg.get_layer(name).output for name in STYLE_LAYERS]\n    content_output = vgg.get_layer(CONTENT_LAYER).output\n    model_outputs = style_outputs + [content_output]\n    return Model(vgg.input, model_outputs)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:37.190804Z","iopub.execute_input":"2024-12-22T16:44:37.191074Z","iopub.status.idle":"2024-12-22T16:44:37.195505Z","shell.execute_reply.started":"2024-12-22T16:44:37.191048Z","shell.execute_reply":"2024-12-22T16:44:37.194620Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this section, we define the functions used to compute the various components of the loss function. The various components are: style loss, content loss, and color loss.\n- Style loss is computed using the Gram matrix, which represents the correlations between the different feature maps and captures the textures and patterns of the style image.\n- Content loss measures the difference between the content of the generated image and the content of the target image, ensuring the generated image retains the important elements of the original content.\n- Color loss ensures that the generated image preserves the colors from the style image, maintaining its visual integrity.\n\nFinally, these individual losses are combined to form the total loss, which will be minimized during training to generate the stylized image.\n\n*This section was generated with the help of ChatGPT.*","metadata":{}},{"cell_type":"code","source":"# Compute Gram matrix\ndef gram_matrix(input_tensor):\n    result = tf.linalg.einsum('bijc,bijd->bcd', input_tensor, input_tensor)\n    input_shape = tf.shape(input_tensor)\n    num_locations = tf.cast(input_shape[1] * input_shape[2], tf.float32)\n    return result / num_locations\n\n# Compute style loss (difference between Gram matrices and the target style)\ndef style_loss(base_style, target_style):\n    S = gram_matrix(base_style)\n    T = gram_matrix(target_style)\n    return tf.reduce_mean(tf.square(S - T))\n\n# Compute content loss\ndef content_loss(base_content, target_content):\n    return tf.reduce_mean(tf.square(base_content - target_content))\n\n# Compute color loss\ndef color_loss(generated_image, content_image):\n    return tf.reduce_mean(tf.square(generated_image - content_image))\n\n# Compute total loss\ndef compute_loss(model, outputs, style_targets, content_targets, generated_image, style_image, content_image):\n    style_outputs = outputs[:len(STYLE_LAYERS)]\n    content_outputs = outputs[len(STYLE_LAYERS):]\n    \n    # Compute style and content losses\n    style_score = tf.add_n([style_loss(style, target) for style, target in zip(style_outputs, style_targets)])\n    content_score = content_loss(content_outputs[0], content_targets[0])\n    style_score *= STYLE_WEIGHT / len(STYLE_LAYERS)\n    content_score *= CONTENT_WEIGHT\n    \n    # Compute color loss\n    color_score = color_loss(generated_image, content_image) * COLOR_WEIGHT\n    \n    # Combine all losses\n    loss = style_score + content_score + color_score\n    return loss","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:37.196646Z","iopub.execute_input":"2024-12-22T16:44:37.197142Z","iopub.status.idle":"2024-12-22T16:44:37.207329Z","shell.execute_reply.started":"2024-12-22T16:44:37.197114Z","shell.execute_reply":"2024-12-22T16:44:37.206535Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this section, we prepare the model for the neural style transfer process. We extract the style and content features from the VGG19 model using the provided style and content images. Next, we initialize the generated image as a copy of the content image and define an Adam optimizer with a learning rate of 0.1, which will be used to minimize the loss during the training process. The optimizer will adjust the generated image to progressively match the style and content targets.","metadata":{}},{"cell_type":"code","source":"# Prepare the model and extract features\nmodel = get_model()\nstyle_targets = model(style_image)[:len(STYLE_LAYERS)]\ncontent_targets = model(content_image)[len(STYLE_LAYERS):]\n\n# Initialize the generated image\ngenerated_image = tf.Variable(content_image, dtype=tf.float32)\ngenerated_img = tf.Variable(content_image, dtype=tf.float32)\n\n# Set up the optimizer\nopt = tf.keras.optimizers.Adam(learning_rate=0.1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:37.210469Z","iopub.execute_input":"2024-12-22T16:44:37.210732Z","iopub.status.idle":"2024-12-22T16:44:39.052781Z","shell.execute_reply.started":"2024-12-22T16:44:37.210708Z","shell.execute_reply":"2024-12-22T16:44:39.052065Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In this section, we define the training step for the neural style transfer process. The train_step function calculates the loss, computes gradients, and updates the generated image using the Adam optimizer. Additionally, the adjust_color_weight function dynamically adjusts the color weight during training to help fine-tune the balance between content and style preservation as the training continues. This ensures that the generated image evolves over time.\n\n*This code section was generated with the help of ChatGPT.*","metadata":{}},{"cell_type":"code","source":"# Training step\n@tf.function\ndef train_step(generated_image):\n    with tf.GradientTape() as tape: # Compute loss and gradients\n        outputs = model(generated_image)\n        loss = compute_loss(model, outputs, style_targets, content_targets, generated_image, style_image, content_image)\n    grad = tape.gradient(loss, generated_image)\n    opt.apply_gradients([(grad, generated_image)]) # Update gradients\n    generated_image.assign(tf.clip_by_value(generated_image, -103.939, 255.0)) # Clip pixel values back to valid range\n    return loss\n\n# Dynamically adjust color weight during training\ndef adjust_color_weight(epoch, max_epochs):\n    new_weight = 1.0 + (2.0 - 1.0) * (epoch / max_epochs)  # Start at 1.0, increase to 2.0\n    COLOR_WEIGHT.assign(new_weight)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:39.053749Z","iopub.execute_input":"2024-12-22T16:44:39.054016Z","iopub.status.idle":"2024-12-22T16:44:39.060283Z","shell.execute_reply.started":"2024-12-22T16:44:39.053990Z","shell.execute_reply":"2024-12-22T16:44:39.059215Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Finally, we run the training loop over a set number of epochs and steps, for each style image. The color weight will be adjusted dynamically. Python will print its progress at regular intervals.\n\n*This code section was build with the help of ChatGPT.*","metadata":{}},{"cell_type":"code","source":"# Training parameters\nepochs_per_style = 10\nsteps_per_epoch = 100\n\n# Function to train for one style image\ndef train_for_style_image(style_image_path, generated_image, content_image):\n    # Load the style image\n    style_image = load_and_process_image(style_image_path)\n    \n    # Extract new style targets for the current style image\n    style_targets = model(style_image)[:len(STYLE_LAYERS)]\n    \n    for epoch in range(epochs_per_style):\n        adjust_color_weight(epoch, epochs_per_style)  # Dynamically adjust COLOR_WEIGHT\n        print(f\"Style: {style_image_path}, Epoch {epoch + 1}, COLOR_WEIGHT: {COLOR_WEIGHT.numpy()}\")\n        \n        for step in range(steps_per_epoch):\n            loss = train_step(generated_image)\n            if step % 10 == 0:\n                print(f\"Epoch {epoch + 1}, Step {step}, Loss: {loss.numpy()}\")\n\n# Main training loop over all style images\nfor i, style_img_path in enumerate(style_image_paths):\n    print(f\"Starting training for style image {i + 1}/{len(style_image_paths)}\")\n    \n    # Train on the current style image\n    train_for_style_image(style_img_path, generated_image, content_image)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:44:39.061697Z","iopub.execute_input":"2024-12-22T16:44:39.062086Z","iopub.status.idle":"2024-12-22T16:48:03.123381Z","shell.execute_reply.started":"2024-12-22T16:44:39.062048Z","shell.execute_reply":"2024-12-22T16:48:03.122685Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Also for only one style image:","metadata":{}},{"cell_type":"code","source":"train_for_style_image('/kaggle/working/train_data/train_1/1642.jpg', generated_img, content_image)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:48:03.124460Z","iopub.execute_input":"2024-12-22T16:48:03.124740Z","iopub.status.idle":"2024-12-22T16:48:46.019995Z","shell.execute_reply.started":"2024-12-22T16:48:03.124713Z","shell.execute_reply":"2024-12-22T16:48:46.019271Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Results <a class=\"anchor\"  id=\"h3\"></a>","metadata":{}},{"cell_type":"markdown","source":"The loss, consisting of the previously mentioned parts, is definetily lower for the network that used 5 style images (loss = ~2.11e7), compared to the loss of the network using 1 style image (loss = ~6.58e7). \nNow, we can convert the final generated image into a displayable format and show it.","metadata":{}},{"cell_type":"code","source":"# Display final results\nfinal_img_1 = deprocess_image(generated_img.numpy())\nplt.imshow(final_img_1)\nplt.axis('off')\nplt.title(\"Final Generated Image, 1 style img\")\nplt.show()\n\nfinal_img = deprocess_image(generated_image.numpy())\nplt.imshow(final_img)\nplt.axis('off')\nplt.title(\"Final Generated Image, 5 style img\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:48:46.021063Z","iopub.execute_input":"2024-12-22T16:48:46.021342Z","iopub.status.idle":"2024-12-22T16:48:46.598865Z","shell.execute_reply.started":"2024-12-22T16:48:46.021315Z","shell.execute_reply":"2024-12-22T16:48:46.597987Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def calculate_ssim(content_image, generated_image):\n    \"\"\"\n    Calculate SSIM (Structural Similarity Index) between the content image and the generated image.\n    \"\"\"\n    # Convert TensorFlow tensors to NumPy arrays if needed\n    if isinstance(generated_image, tf.Tensor) or isinstance(generated_image, tf.Variable):\n        generated_image = generated_image.numpy()\n\n    # Remove batch dimension\n    content_image = content_image[0]\n    generated_image = generated_image[0]\n\n    # Deprocess images\n    content_img = deprocess_image(content_image)\n    generated_img = deprocess_image(generated_image)\n\n    # Compute SSIM with explicit window size\n    ssim = compare_ssim(content_img, generated_img, win_size=7, channel_axis=2)  # Specify win_size and channel_axis\n    return ssim","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:48:46.599955Z","iopub.execute_input":"2024-12-22T16:48:46.600226Z","iopub.status.idle":"2024-12-22T16:48:46.605773Z","shell.execute_reply.started":"2024-12-22T16:48:46.600200Z","shell.execute_reply":"2024-12-22T16:48:46.604874Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Evaluate SSIM after training\nssim_score1 = calculate_ssim(content_image, generated_img)\nssim_score = calculate_ssim(content_image, generated_image)\n\n# Print  SSIM score\nprint(f\"SSIM (Content Similarity) using 1 style image: {ssim_score1}\")\nprint(f\"SSIM (Content Similarity) using 5 style images: {ssim_score}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-22T16:48:46.606852Z","iopub.execute_input":"2024-12-22T16:48:46.607106Z","iopub.status.idle":"2024-12-22T16:48:46.662417Z","shell.execute_reply.started":"2024-12-22T16:48:46.607082Z","shell.execute_reply":"2024-12-22T16:48:46.661605Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Our main results are based on the visual evaluation of the generated images, supported by the Structural Similarity Index (SSIM) as an additional metric. SSIM, developed by Wang et al. (2004), models image distortion by assessing loss correlation, luminance distortion, and contrast distortion. Unlike Mean Squared Error (MSE), which primarily focuses on pixel-wise differences, SSIM has been shown to better capture perceptual quality making it a valuable addition to our project.\n\nSSIM values range from -1 to 1, where higher values indicate greater structural similarity between the content and generated images. For our research, an SSIM score of approximately 0.35 reflects a reasonable balance. Given that the style applied is Impressionism, characterized by distinct brushstrokes, and the content image is a photograph with no such strokes, a lower SSIM is expected and even desirable, as long as it does not approach negative SSIM. The latter is important when inspecting this notebook, since the SSIM of the 1 style image network is higher. This indicates successful stylization while retaining some structural consistency with the original content.","metadata":{}},{"cell_type":"markdown","source":"# 4. Conclusion/ Discussion <a class=\"anchor\"  id=\"h4\"></a>","metadata":{}},{"cell_type":"markdown","source":"## 4.1 Generalizability <a class=\"anchor\"  id=\"41\"></a>\n\nIn our small project, we are very content with the impressionistic dog painting we have created. Next to that, we can see a definite improved from the 1 style image network to the 5 style images network. However, due to the time & memory limitations we couldn't investigate how different colors & other pictures (like black and white) would transfer.\n\nCould our model generalize to more complex styles like Abstract art, Surrealism, or Cubism? How would our dog do in the complex style of Abstract Art, or even an artists own style, like Van Gogh, Mondriaan etc.","metadata":{}},{"cell_type":"markdown","source":"## 4.2 Opportunities for Improvement <a class=\"anchor\"  id=\"42\"></a>\n\nOur initial goal was to let the user choose between 4 styles for their image conversion: Impressionism, Realism, Romanticism & Expressionism. However, due to Kaggle's memory limitations, we were constrained in our data options and model complexity, which forced us to alter the direction of our project. This leaves room for improvement. Especially, when there is access to bigger memory bases.\n\nAnother idea in order to make this project more practical would be to make an application, when more memory is available. Finetuning different styles & artistic directions so people could insert own pictures into the app, having the style transfer done through an external server to save computation time. We could use user feedback as part of the development process, allowing adjustments of style transfer results based on individual preferences. ","metadata":{}},{"cell_type":"markdown","source":"## 4.3 Important Takeaways for Practitioners <a class=\"anchor\"  id=\"43\"></a>\nUsing the VGG19 pre-trained model and the Painter by Numbers dataset saved development time in our neural style transfer project. User-centric design proved crucial, as adjusting weights was time-intensive. An interactive interface for dynamic parameter tuning could simplify the process, enhancing both efficiency and model performance.","metadata":{}},{"cell_type":"markdown","source":"## 4.4 Important Takeaways for Researchers <a class=\"anchor\"  id=\"44\"></a>\nFuture work could combine styles, like Impressionism and Abstract art, to study interactions. Addressing computational limits, we recommend efficient projects using pre-trained models, compressed images, and tools like sliders for post-transfer tuning, balancing resource use with meaningful, innovative results.","metadata":{}},{"cell_type":"markdown","source":"# 5. References <a class=\"anchor\"  id=\"h5\"></a>\n\nfchollet. (2020). Neural style transfer. Keras. https://keras.io/examples/generative/neural_style_transfer/\n\nPainting, P. (2022). The 10 Traits of Impressionism - PopUp Painting - Medium. Medium. https://medium.com/@PopUpPainting/the-10-traits-of-impressionism-2a2c045795c7\n\nTelega, S., PhD. (2024). Neural Style Transfer in Keras — step by step. Part I — Loss function for Content and Style. Medium. https://medium.com/@telega.slawomir.ai/neural-style-transfer-in-keras-step-by-step-part-i-loss-function-for-content-and-style-85227b7f3586\n\nXu, Y. (2024). CNN-based image style transformation--Using VGG19. Applied And Computational Engineering, 39(1), 130–136. https://doi.org/10.54254/2755-2721/39/20230589\n\nZ. Wang, A. C. Bovik, H. R. Sheikh and E. P. Simoncelli, \"Image quality assessment: from error visibility to structural similarity\", IEEE Transactions on Image Processing, vol. 13, no. 4, pp. 600-612, 2004.","metadata":{}},{"cell_type":"markdown","source":"# 6. Division of Labour <a class=\"anchor\"  id=\"h6\"></a>\n\nLoading and preprocessing images: Philip  \nMake a basis for the neural transfer: Charlotte  \nFinetuning the style transfer model: Ralph and Philip  \nCreating the poster: Ralph  \nFinalizing the notebook, adding theory: Charlotte, Philip and Ralph","metadata":{}}]}