{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":113558,"databundleVersionId":14456136,"sourceType":"competition"}],"dockerImageVersionId":31192,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this competition, our goal is to identify manipulated or fraudulent images that appear in scientific papers and journals. Traditionally, this task is performed by expert reviewers who rely on years of experience to spot inconsistencies or abnormalities. However, this manual process is slow, tedious, and impossible to scale across the rapidly growing volume of publications.\n\nTo overcome these limitations, we aim to build an automated machine learning model capable of determining whether a given image is authentic or fake, enabling faster and more reliable scientific integrity checks at scale.","metadata":{}},{"cell_type":"code","source":"import pathlib\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt \nimport cv2\nimport random","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:19.438306Z","iopub.execute_input":"2025-12-03T16:34:19.438579Z","iopub.status.idle":"2025-12-03T16:34:21.756830Z","shell.execute_reply.started":"2025-12-03T16:34:19.438550Z","shell.execute_reply":"2025-12-03T16:34:21.755939Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"BASE_DIR = pathlib.Path(\"/kaggle/input/recodai-luc-scientific-image-forgery-detection\")\nTRAIN_IMG_DIR = BASE_DIR / \"train_images\"\nTRAIN_MASK_DIR = BASE_DIR / 'train_masks'\n\nTEST_IMG_DIR = BASE_DIR/ 'test_images'\n\nSUPPLE_IMG_DIR = BASE_DIR / 'supplemental_images'\nSUPPLE_MASK_DIR = BASE_DIR / 'supplemental_masks'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:21.758357Z","iopub.execute_input":"2025-12-03T16:34:21.758773Z","iopub.status.idle":"2025-12-03T16:34:21.763338Z","shell.execute_reply.started":"2025-12-03T16:34:21.758741Z","shell.execute_reply":"2025-12-03T16:34:21.762449Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('train directory overview')\nfor i in TRAIN_IMG_DIR.iterdir():\n    print(i.stem, len(list(i.iterdir())))\n\n\nprint('train masks', len(list(TRAIN_MASK_DIR.iterdir())))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:21.764465Z","iopub.execute_input":"2025-12-03T16:34:21.764771Z","iopub.status.idle":"2025-12-03T16:34:21.968578Z","shell.execute_reply.started":"2025-12-03T16:34:21.764742Z","shell.execute_reply":"2025-12-03T16:34:21.966575Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('test directory overview ')\nprint(TEST_IMG_DIR.stem, len(list(TEST_IMG_DIR.iterdir())))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:21.970318Z","iopub.execute_input":"2025-12-03T16:34:21.974188Z","iopub.status.idle":"2025-12-03T16:34:21.983969Z","shell.execute_reply.started":"2025-12-03T16:34:21.974128Z","shell.execute_reply":"2025-12-03T16:34:21.983180Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print('supplemental directory overview')\nprint(SUPPLE_IMG_DIR.stem, len(list(SUPPLE_IMG_DIR.iterdir())))\nprint(SUPPLE_MASK_DIR.stem, len(list(SUPPLE_MASK_DIR.iterdir())))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:21.984581Z","iopub.execute_input":"2025-12-03T16:34:21.984833Z","iopub.status.idle":"2025-12-03T16:34:22.021041Z","shell.execute_reply.started":"2025-12-03T16:34:21.984804Z","shell.execute_reply":"2025-12-03T16:34:22.020293Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n### Dataset Structure & Key Observations\n\n**Training Data**\nThe training dataset contains two categories of images:\n\nAuthentic – genuine, unaltered images\n\nForged – manipulated images, each accompanied by a mask highlighting the tampered region\n\n\n\n🔍 **Test Data**\n\nThe test directory contains only a single image, which the model must classify as authentic or forged.\n\nIf the image is forged, the model must also produce the corresponding mask.\n\n\n🗂 **Supplementary Data**\n\nThe supplementary directory includes additional images, all of which come with masks.\n\nThe dataset description does not explicitly state whether these images are authentic or forged.\n\n\n📌 **Assumption**\n\nSince every supplementary image includes a mask, we assume that these images are also forged and can be incorporated into the forged training set.","metadata":{}},{"cell_type":"code","source":"def show_img(img_dir,  mask_dir, img_label, only_img=False, SAMPLES_SIZE=5):\n    for i in range(SAMPLES_SIZE):\n        file_path = random.choice(list((img_dir).iterdir()))\n        img = cv2.imread(file_path)\n    \n        plt.figure(figsize=(12, 12))\n        plt.subplot(1, 2, 1)\n        plt.imshow(img)\n        plt.title(img_label)\n\n        if(not only_img):\n            item_id = file_path.stem\n            mask_path = mask_dir / f\"{item_id}.npy\"\n        \n            mask = np.load(mask_path)\n            mask = np.array(mask, dtype=np.float32)[0]\n            plt.subplot(1,2,2)\n            plt.imshow(mask, cmap=\"gray\")\n            plt.title(\"Mask\")\n\n        plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:22.022050Z","iopub.execute_input":"2025-12-03T16:34:22.022264Z","iopub.status.idle":"2025-12-03T16:34:22.028816Z","shell.execute_reply.started":"2025-12-03T16:34:22.022246Z","shell.execute_reply":"2025-12-03T16:34:22.027952Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# show authentic images in train directory\n\nshow_img(TRAIN_IMG_DIR / 'authentic', ' ', \"Authentic train image\", True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:22.029742Z","iopub.execute_input":"2025-12-03T16:34:22.030082Z","iopub.status.idle":"2025-12-03T16:34:23.387515Z","shell.execute_reply.started":"2025-12-03T16:34:22.030026Z","shell.execute_reply":"2025-12-03T16:34:23.386651Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# show image in test directory\n\nshow_img(TEST_IMG_DIR, ' ', \"Test Image\" , True, 1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:23.388397Z","iopub.execute_input":"2025-12-03T16:34:23.388639Z","iopub.status.idle":"2025-12-03T16:34:23.947183Z","shell.execute_reply.started":"2025-12-03T16:34:23.388617Z","shell.execute_reply":"2025-12-03T16:34:23.946285Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# show forged images in train directory\n\nshow_img(TRAIN_IMG_DIR / \"forged\", TRAIN_MASK_DIR, \"Forged Image\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:23.948084Z","iopub.execute_input":"2025-12-03T16:34:23.948316Z","iopub.status.idle":"2025-12-03T16:34:26.820825Z","shell.execute_reply.started":"2025-12-03T16:34:23.948297Z","shell.execute_reply":"2025-12-03T16:34:26.819944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# show images in Supplementry directory\n\nshow_img(SUPPLE_IMG_DIR, SUPPLE_MASK_DIR, \"Supplement Image\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-12-03T16:34:26.822642Z","iopub.execute_input":"2025-12-03T16:34:26.822912Z","iopub.status.idle":"2025-12-03T16:34:35.376057Z","shell.execute_reply.started":"2025-12-03T16:34:26.822868Z","shell.execute_reply":"2025-12-03T16:34:35.375245Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Obervation**\n\n\n- It's just an observation that, images avaialble in the dataset are belong to various scientific domains. It contains various patterns, visuals, textures. Hence we need a strong generalize model to capture the finnest details which can distinguish between authentic and forged.\n\n- The dataset also includes images of varying shapes and resolutions, which means preprocessing steps such as resizing, normalization, and aspect-ratio handling will be essential for consistent model performance.\n","metadata":{}},{"cell_type":"markdown","source":"**Approach**\n\nThe approach i'll be using is 2 phase approach.\n\n1) **Image Classification** \n    In the Phase 1, the image get's classified into 1 category Authentic or Forged by a      classification model. If the image is Authentic and no processing.\n   \n2) **Mask Prediction (Segmentation)** \n    for the Image which is label forged, we will find it's mask with other model.","metadata":{}},{"cell_type":"markdown","source":"link of phase 1 demonstration: https://www.kaggle.com/code/tany1404/recod-ai-luc-base-model-phase-1-v0","metadata":{}}]}