{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":98450,"databundleVersionId":11749951,"sourceType":"competition"}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I didnt want to start a perfect cause i was unsure what the invalid data was, i tried to stick with CNN and neural networks and went in open minded the first few tries to atleast have a starting point which puts me on the map of progress, so when i know where i want to go, i will know where i am, which makes things so much easier, everything went as plan at first, i identified the invalid data, made data generators and filters, optimized some tricks to the model and decided to continue with (3D) CNN, and just as i thought everything was going fine and according to plan, little did i know that then was the time everything was about to go wrong, just as i progressivly made my model better, i encountered a hell lot of memory errors, i run everything on cpu no gpu, and everything was a little difficult to digest, what i would say my strategy is that i minimized the shape for digestion then changed it at last for the target shape, which simplified everything, so to conclude, T1 - T4 starting out and knowing the errors and everything. T5 - T6 Memory issues, T7 - T9, tried on a device with a gpu but didnt trigger gpu so a lot of time wasted there, T10 V1 and V2 were made to run on strong gpus but since the chance to use the gpu again was not available then, and that was my last wall, i tried to play around with the code and made the T10 V2 Light, optimized for my cpu using the strat of the shape and so. It yielded the best results, but there were like 50+ tries after but i didnt want to make more python programs, i just had 3 python programs open and done everything in them, for instance, i typed the code in first python program, made modifications, if bad, reset to how the code was before editing it, if good, move it to the second python file to continue modifications, the cycle keeps going till never ending, if file 3 was good and needed modifications, i replace the first python program with it, and so. But throughout all these tries the T10 V2 Light had the best results, the difference in private score brought to my mind that there could possibly be models not chosen due to low public score, could have possibly had really good private score, but i am still studing that matter, I would say that it is my prayers that got me chosen in the private, other than that, the code was somewhat good as well, i am still seeking to improve. So in conclusion, the \"strategy\" i used is making the light version, cause the tactic is minimizing the shaping to digest fast to swiftly train the model and at the end (of the training python file) yield the target shape of (32, 32, 32), and in the prediction python file, yield the shape of (128, 128, 125). I am really open to any questions, i hope my answer was enough for the requirement!! :)","metadata":{}},{"cell_type":"code","source":"import kagglehub\n#kagglehub.login()\nbeyond_visible_spectrum_ai_for_agriculture_2025_path = kagglehub.competition_download('beyond-visible-spectrum-ai-for-agriculture-2025')\nprint('Data source import complete.')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\nThe First python program is the main one for data processing and training the model plus saving. Keep in mind that the one that worked had 10 epochs and early stop patience set to 3 XD, I massivly updated them in attempts to get better results also when I tried batch sizes of 2, 4, 8, 16, 32, and 64.","metadata":{}},{"cell_type":"markdown","source":"I minimized epochs to 10, and patience to 3 which was the very begining of which i recall, feel free to increase them (or run several times) to get good results!","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models # type: ignore\nfrom tensorflow.keras.models import Model # type: ignore\nfrom tensorflow.keras.optimizers import Adam # type: ignore\nfrom tensorflow.keras.callbacks import ModelCheckpoint, EarlyStopping, ReduceLROnPlateau # type: ignore\nfrom tensorflow.keras.utils import Sequence # type: ignore\nfrom sklearn.model_selection import train_test_split\nfrom tqdm import tqdm\nimport random\nimport cv2\n\n# Was hoping to use my relative's pc with GPU but the chance didnt come, i just left this as an if/else statement in case i ever came across it\n\ngpus = tf.config.list_physical_devices('GPU')\nif gpus:\n    print(f\"GPU detected: {gpus[0].name}\")\n    tf.config.experimental.set_memory_growth(gpus[0], True)\nelse:\n    print(\"No GPU detected. Using CPU.\")\n\nnpy_folder = \"/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot\"\ncsv_path = \"/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/train.csv\"\ntarget_shape = (32, 32, 32)\nbatch_size = 1\nepochs = 15\nmodel_path = \"mini_model_cpu3.keras\"\n\nclass DataGenerator(Sequence):\n    def __init__(self, file_paths, labels, batch_size, target_shape, npy_folder, augment=False):\n        self.raw_file_paths = file_paths\n        self.raw_labels = labels\n        self.batch_size = batch_size\n        self.target_shape = target_shape\n        self.npy_folder = npy_folder\n        self.augment = augment\n        self.file_paths, self.labels = self._filter_valid_files() # Filter (modify later)\n        self.indexes = np.arange(len(self.file_paths))\n\n    def _filter_valid_files(self):\n        valid_paths = []\n        valid_labels = []\n        for path, label in tqdm(zip(self.raw_file_paths, self.raw_labels), desc=\"Filtering corrupt files\"):\n            try:\n                full_path = os.path.join(self.npy_folder, path.strip())\n                data = np.load(full_path)\n                if data.size == 0 or len(data.shape) != 3:\n                    raise ValueError(\"Empty or malformed data\")\n                valid_paths.append(path)\n                valid_labels.append(label)\n            except Exception as e:\n                print(f\"Skipping file: {path} | Reason: {e}\")\n        return valid_paths, valid_labels\n\n    def __len__(self):\n        return int(np.floor(len(self.file_paths) / self.batch_size))\n\n    def __getitem__(self, index):\n        batch_indexes = self.indexes[index * self.batch_size:(index + 1) * self.batch_size]\n        batch_paths = [self.file_paths[k] for k in batch_indexes]\n        batch_labels = [self.labels[k] for k in batch_indexes]\n\n        X_batch = np.zeros((self.batch_size, *self.target_shape, 1), dtype=np.float32)\n        for i, file_path in enumerate(batch_paths):\n            try:\n                full_path = os.path.join(self.npy_folder, file_path.strip())\n                data = np.load(full_path)\n                data = np.expand_dims(data, axis=-1)\n                resized = np.zeros((*self.target_shape, 1), dtype=np.float32)\n                for z in range(min(data.shape[2], self.target_shape[2])):\n                    resized[:, :, z, 0] = cv2.resize(data[:, :, z], self.target_shape[:2])\n                X_batch[i] = resized\n            except Exception as e:\n                print(f\"Error batch loading: {file_path} | Reason: {e}\")\n                continue \n\n        return X_batch, np.array(batch_labels)\n\n    def on_epoch_end(self):\n        np.random.shuffle(self.indexes)\n        \ndef build_mini_model(input_shape=(32, 32, 32, 1)): # CNN of shape (32, 32, 32, 1) # Attempt 27 (afterupdate)\n    inputs = layers.Input(shape=input_shape)\n\n    x = layers.Conv3D(16, (3, 3, 3), padding='same', activation='relu')(inputs)\n    x = layers.MaxPooling3D((2, 2, 2))(x)\n\n    x = layers.Conv3D(32, (3, 3, 3), padding='same', activation='relu')(x)\n    x = layers.MaxPooling3D((2, 2, 2))(x)\n\n    x = layers.Conv3D(64, (3, 3, 3), padding='same', activation='relu')(x)\n    x = layers.GlobalAveragePooling3D()(x)\n\n    x = layers.Dense(64, activation='relu')(x)\n    x = layers.Dropout(0.3)(x)\n    output = layers.Dense(1, activation='linear')(x)\n\n    model = Model(inputs, output)\n    return model\n\ndf = pd.read_csv(csv_path)\nfile_paths = df['id'].values\nlabels = df['label'].values\n\nX_train, X_val, y_train, y_val = train_test_split(file_paths, labels, test_size=0.2, random_state=42)\n\ntrain_gen = DataGenerator(X_train, y_train, batch_size, target_shape, npy_folder)\nval_gen = DataGenerator(X_val, y_val, batch_size, target_shape, npy_folder)\n\nmodel = build_mini_model(input_shape=(*target_shape, 1))\nmodel.compile(optimizer=Adam(0.001), loss='mse', metrics=['mae'])\n\ncallbacks = [\n    EarlyStopping(patience=3, restore_best_weights=True, verbose=1),\n    ReduceLROnPlateau(factor=0.5, patience=2, min_lr=1e-6, verbose=1),\n    ModelCheckpoint(model_path, save_best_only=True, monitor='val_loss', mode='min', verbose=1)\n]\n\nmodel.fit(train_gen, validation_data=val_gen, epochs=epochs, callbacks=callbacks, verbose=1)\n\n# Hey There! I tried a lot of modifications to this like batch size of 2, 4 and 8, i modified some stuff but the current state has the best ones, batch size 1, it is also very very cpu friendly, this model is my best of all!!!\n\n#There is another python program I use for prediction, i will send it as well!!","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The next python program is used to convert the (32, 32, 32) from the primary python program to the (128, 128, 125) needed for the challenge as the main shape to predict upon!","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nfrom tensorflow.keras.models import load_model  # type: ignore\n\n\nnpy_folder = \"/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/ot/ot\"\ncsv_path = \"/kaggle/input/beyond-visible-spectrum-ai-for-agriculture-2025/test.csv\"\nmodel_path = \"mini_model_cpu3.keras\"\ntarget_shape = (128, 128, 125)\nmodel_input_shape = (32, 32, 32)\n\nif not os.path.isfile(model_path):\n    raise FileNotFoundError(f\"file not found: {model_path}\")\nmodel = load_model(model_path)\nprint(\"Model loaded\")\n\ndef preprocess_npy(file_path, target_shape=(128, 128, 125), model_input_shape=(32, 32, 32)):\n    try:\n        data = np.load(file_path)\n\n        pad_width = [(0, max(0, target_shape[i] - data.shape[i])) for i in range(3)]\n        data = np.pad(data, pad_width, mode='constant')\n\n        slices = tuple(slice(0, target_shape[i]) for i in range(3))\n        data = data[slices]\n\n        data = data[:model_input_shape[0], :model_input_shape[1], :model_input_shape[2]]\n\n        data = data.astype(np.float32)\n        max_val = np.max(data)\n        if max_val > 0:\n            data /= max_val\n\n        data = data[np.newaxis, ..., np.newaxis]\n        return data\n\n    except Exception as e:\n        print(f\"Failed or process {file_path}: {e}\")\n        return None\n    \ndef safe_predict(file_path, target_shape, model_input_shape, model):\n    try:\n        volume = preprocess_npy(file_path, target_shape, model_input_shape)\n        if volume is None:\n            return None\n        return model.predict(volume, verbose=0)[0][0]\n    except Exception as e:\n        print(f\"Failed to process {file_path}: {e}\")\n        return None\n    \ndf_test = pd.read_csv(csv_path)\ndf_test['id'] = df_test['id'].astype(str).str.replace('.npy', '', regex=False).str.strip()\n\npredictions = []\nskipped_indices = []\n\nfor idx, file_id in enumerate(tqdm(df_test['id'], desc=\"Processing test samples\")):\n    file_path = os.path.join(npy_folder, f\"{file_id}.npy\")\n    \n    pred = safe_predict(file_path, target_shape, model_input_shape, model)\n    \n    if pred is None:\n        predictions.append(None)\n        skipped_indices.append(idx)\n        continue\n\n    predictions.append(pred)\n    \nnum_skipped = len(skipped_indices)\nif num_skipped > 0:\n    print( (num_skipped) + \" samples failed to process. Filling with mean prediction.\")\n    valid_preds = [p for p in predictions if p is not None]\n    fallback_value = np.mean(valid_preds) if valid_preds else 0.0\n    for i in skipped_indices:\n        predictions[i] = fallback_value\n        \ndf_test['label'] = predictions\ndf_test.to_csv(\"predictions.csv\", index=False)\nprint(\"Predictions saved Finnaly!\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"And there you have it! That entire thing can run on literally ANYTHING!! No matter what specs you have, i got my results from running this on just a CPU, imagine how much better you could get? If you wanna try and even get better results, try playing a bit around with the shape and run multiple times because having a batch size of 1 somewhat makes everything random, Thank you!!","metadata":{}}]}