{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":11000,"databundleVersionId":875412}],"dockerImageVersionId":31328,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# LANL Earthquake Prediction \n### Hazırlayan: Serdar ÖNAL (Civil Engineer)\n\n## 📌 Proje Amacı\nBu çalışma, laboratuvar ortamında simüle edilen sismik verileri (acoustic data) kullanarak, bir sonraki depremin ne zaman gerçekleşeceğini (time to failure) tahmin etmeyi amaçlar. Bir inşaat mühendisi perspektifiyle, yapı sağlığı izleme sistemlerine temel teşkil edebilecek bir sismik analiz modelidir.\n\n## 🛠️ Metodoloji ve Özellik Mühendisliği\nHam veri, her biri 150.000 satırdan oluşan sinyal bloklarından oluşmaktadır. Modelin bu veriyi anlamlandırabilmesi için \"yazıya dökme\" (feature engineering) işlemi uygulanmıştır:\n\n*   **Temel İstatistikler:** Ortalama, standart sapma, minimum ve maksimum değerler.\n*   **Doku ve Oynaklık (Volatility):** Sinyalin standart sapması (dalgalanma), deprem öncesi artan mikro-çatlak seslerini yakalamak için kullanılmıştır.\n*   **Çarpıklık (Skewness):** Sinyal dağılımının asimetrisini ölçerek, ani enerji boşalımları (spike) tespit edilmiştir.\n\n## 🧠 Model Mimarisi\nProjede yüksek performanslı bir gradyan artırma algoritması olan **CatBoost Regressor** tercih edilmiştir:\n*   **Strateji:** Gürültülü (noisy) sismik verilerde aşırı öğrenmeyi (overfitting) engelleyen \"Overfitting Detector\" aktif edilmiştir.\n*   **Kayıp Fonksiyonu:** MAE (Mean Absolute Error).","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport kagglehub\nimport os","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:13:30.915084Z","iopub.execute_input":"2026-05-12T23:13:30.915834Z","iopub.status.idle":"2026-05-12T23:13:31.914836Z","shell.execute_reply.started":"2026-05-12T23:13:30.915801Z","shell.execute_reply":"2026-05-12T23:13:31.913862Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Veriyi İndirme\nprint(\"⏳ Veri indiriliyor, lütfen bekleyin...\")\npath = kagglehub.competition_download('LANL-Earthquake-Prediction')\nprint(\"✅ Veri yolu:\", path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T22:51:59.736862Z","iopub.execute_input":"2026-05-12T22:51:59.737161Z","iopub.status.idle":"2026-05-12T22:55:30.408913Z","shell.execute_reply.started":"2026-05-12T22:51:59.737136Z","shell.execute_reply":"2026-05-12T22:55:30.407873Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Dosya Kontrolü\nfiles = os.listdir(path)\nprint(f\"📁 Klasördeki dosyalar: {files}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T22:55:44.466598Z","iopub.execute_input":"2026-05-12T22:55:44.467494Z","iopub.status.idle":"2026-05-12T22:55:44.475775Z","shell.execute_reply.started":"2026-05-12T22:55:44.467460Z","shell.execute_reply":"2026-05-12T22:55:44.474809Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Veriyi Parçalı Okuma (Chunking)\n# İlk 10 milyon satır\ntrain_sample = pd.read_csv(\n    os.path.join(path, 'train.csv'), \n    nrows=10000000, \n    dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32}\n)\n\nprint(\"\\n🔍 İlk 10 Milyon Satırın Özeti:\")\nprint(train_sample.head())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T22:55:49.221392Z","iopub.execute_input":"2026-05-12T22:55:49.222239Z","iopub.status.idle":"2026-05-12T22:55:51.587215Z","shell.execute_reply.started":"2026-05-12T22:55:49.222205Z","shell.execute_reply":"2026-05-12T22:55:51.586393Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Akustik Veri (Sismik Sinyal)\nfig, ax1 = plt.subplots(figsize=(12, 6))\nax1.set_xlabel('Zaman (Örnek sayısı)')\nax1.set_ylabel('Akustik Sinyal', color='tab:blue')\nax1.plot(train_sample['acoustic_data'].values, color='tab:blue', alpha=0.6)\nax1.tick_params(axis='y', labelcolor='tab:blue')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T22:56:26.541411Z","iopub.execute_input":"2026-05-12T22:56:26.542385Z","iopub.status.idle":"2026-05-12T22:56:28.913044Z","shell.execute_reply.started":"2026-05-12T22:56:26.542346Z","shell.execute_reply":"2026-05-12T22:56:28.912196Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Time to Failure (Depreme Kalan Süre)\nfig, ax1 = plt.subplots(figsize=(12, 6))\nax2 = ax1.twinx()\nax2.set_ylabel('Depreme Kalan Süre (sn)', color='tab:red')\nax2.plot(train_sample['time_to_failure'].values, color='tab:red', linewidth=3)\nax2.tick_params(axis='y', labelcolor='tab:red')\n\nplt.title('Akustik Sinyal ve Hedef Değişken (İlk 10M Satır)')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T22:57:14.451906Z","iopub.execute_input":"2026-05-12T22:57:14.452206Z","iopub.status.idle":"2026-05-12T22:57:16.027942Z","shell.execute_reply.started":"2026-05-12T22:57:14.452180Z","shell.execute_reply":"2026-05-12T22:57:16.027006Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Verinin farklı bölgelerden örnekler alma (Ram krizi yaşamamak için )\n# İlk 20 milyon satırı inceleme\ntrain_subset = pd.read_csv(\n    os.path.join(path, 'train.csv'), \n    nrows=20000000, \n    dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32}\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:15:32.753036Z","iopub.execute_input":"2026-05-12T23:15:32.753755Z","iopub.status.idle":"2026-05-12T23:15:36.850658Z","shell.execute_reply.started":"2026-05-12T23:15:32.753721Z","shell.execute_reply":"2026-05-12T23:15:36.849659Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Sinyal Dağılım Analizi (Histogram)\nplt.figure(figsize=(12, 5))\nsns.histplot(train_subset['acoustic_data'], bins=100, kde=True, color='forestgreen')\nplt.title('Akustik Sinyal Dağılımı (Genlik Analizi)')\nplt.xlabel('Sinyal Genliği')\nplt.ylabel('Frekans')\nplt.grid(True, alpha=0.3)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:15:40.035613Z","iopub.execute_input":"2026-05-12T23:15:40.036469Z","iopub.status.idle":"2026-05-12T23:17:05.430141Z","shell.execute_reply.started":"2026-05-12T23:15:40.036432Z","shell.execute_reply":"2026-05-12T23:17:05.429130Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Birden Fazla Deprem Döngüsünü İnceleme\n# Verideki 'time_to_failure' değerinin sıfıra indirme ve tekrar sıçradığı anları yakalama\nfig, ax1 = plt.subplots(figsize=(15, 6))\n\nax1.set_xlabel('Index')\nax1.set_ylabel('Akustik Sinyal', color='gray')\nax1.plot(train_subset['acoustic_data'], color='gray', alpha=0.5, label='Akustik Veri')\nax1.tick_params(axis='y', labelcolor='gray')\n\nax2 = ax1.twinx()\nax2.set_ylabel('Depreme Kalan Süre (sn)', color='red')\nax2.plot(train_subset['time_to_failure'], color='red', linewidth=2, label='Time to Failure')\nax2.tick_params(axis='y', labelcolor='red')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:17:36.586321Z","iopub.execute_input":"2026-05-12T23:17:36.587106Z","iopub.status.idle":"2026-05-12T23:17:45.139229Z","shell.execute_reply.started":"2026-05-12T23:17:36.587070Z","shell.execute_reply":"2026-05-12T23:17:45.138301Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Feature Processing\nfrom scipy.stats import skew, kurtosis\nfrom tqdm import tqdm\nfrom scipy.fftpack import fft","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:36:42.623873Z","iopub.execute_input":"2026-05-12T23:36:42.624341Z","iopub.status.idle":"2026-05-12T23:36:42.661725Z","shell.execute_reply.started":"2026-05-12T23:36:42.624256Z","shell.execute_reply":"2026-05-12T23:36:42.660736Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"rows = 150000\n\n# Gelişmiş Özellik Mühendisliği Fonksiyonu\ndef create_features_advanced(seg):\n    X = pd.DataFrame()\n    xc = seg['acoustic_data'].values\n    \n    # --- Zaman Domaini Özellikleri ---\n    X['mean'] = [xc.mean()]\n    X['std'] = [xc.std()]\n    X['max'] = [xc.max()]\n    X['min'] = [xc.min()]\n    X['abs_max'] = [np.abs(xc).max()]\n    X['skew'] = [skew(xc)]\n    X['kurtosis'] = [kurtosis(xc)]\n    \n    # --- Frekans Domaini (FFT) Özellikleri ---\n    # Sinyali frekans bileşenlerine ayırma\n    zc = np.fft.fft(xc)\n    realFFT = np.real(zc)\n    imagFFT = np.imag(zc)\n    \n    X['Rmean'] = [realFFT.mean()]\n    X['Rstd'] = [realFFT.std()]\n    X['Imean'] = [imagFFT.mean()]\n    X['Istd'] = [imagFFT.std()]\n    \n    \n    return X","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:39:13.507692Z","iopub.execute_input":"2026-05-12T23:39:13.508394Z","iopub.status.idle":"2026-05-12T23:39:13.516643Z","shell.execute_reply.started":"2026-05-12T23:39:13.508361Z","shell.execute_reply":"2026-05-12T23:39:13.515789Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_reader = pd.read_csv(os.path.join(path, 'train.csv'), \n                           chunksize=rows, \n                           dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n\nX_train = []\ny_train = []\n\nfor chunk in tqdm(train_reader):\n    # Özellikleri hesapla\n    features = create_features_advanced(chunk)\n    X_train.append(features)\n    \n    # Hedef değişken (her bloğun sonundaki deprem zamanı)\n    y_train.append(chunk['time_to_failure'].values[-1])\n\n# DataFrame'leri birleştir\nX_train = pd.concat(X_train).reset_index(drop=True)\ny_train = pd.Series(y_train)\n\nprint(f\"✅ İşlem Tamam! {X_train.shape[0]} segment hazır.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:40:13.792931Z","iopub.execute_input":"2026-05-12T23:40:13.793934Z","iopub.status.idle":"2026-05-12T23:43:43.599273Z","shell.execute_reply.started":"2026-05-12T23:40:13.793895Z","shell.execute_reply":"2026-05-12T23:43:43.598355Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Korelasyon Analizi\n\n# Özellikler ve Hedef Değişkeni birleştirme\ntrain_analysis = X_train.copy()\ntrain_analysis['target'] = y_train.values\n\n# Korelasyon matrisi\nplt.figure(figsize=(12, 8))\ncorrelations = train_analysis.corr()['target'].sort_values(ascending=False)\ncorrelations.plot(kind='barh', color='skyblue')\nplt.title('Özelliklerin Deprem Zamanı (Target) ile Korelasyonu')\nplt.xlabel('Korelasyon Katsayısı')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:46:21.062813Z","iopub.execute_input":"2026-05-12T23:46:21.063653Z","iopub.status.idle":"2026-05-12T23:46:21.251963Z","shell.execute_reply.started":"2026-05-12T23:46:21.063618Z","shell.execute_reply":"2026-05-12T23:46:21.251056Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Model Eğitimine Geçiş: LightGBM\n\nimport lightgbm as lgb\nfrom sklearn.model_selection import KFold\nfrom sklearn.metrics import mean_absolute_error\n\n# Çapraz Doğrulama (Cross-Validation) Ayarları\nfolds = KFold(n_splits=5, shuffle=True, random_state=42)\n\nparams = {\n    'num_leaves': 31,\n    'min_data_in_leaf': 10,\n    'objective': 'regression',\n    'max_depth': -1,\n    'learning_rate': 0.01,\n    \"boosting\": \"gbdt\",\n    \"feature_fraction\": 0.9,\n    \"bagging_freq\": 1,\n    \"bagging_fraction\": 0.9,\n    \"metric\": 'mae',\n    \"lambda_l1\": 0.1,\n    \"verbosity\": -1,\n    \"nthread\": -1,\n    \"random_state\": 4590\n}\n\noof_predictions = np.zeros(len(X_train))\nfeature_importance_df = pd.DataFrame()\n\nprint(\" Model Eğitimi Başlıyor...\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:49:10.752389Z","iopub.execute_input":"2026-05-12T23:49:10.753432Z","iopub.status.idle":"2026-05-12T23:49:14.397677Z","shell.execute_reply.started":"2026-05-12T23:49:10.753393Z","shell.execute_reply":"2026-05-12T23:49:14.396810Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for fold_, (trn_idx, val_idx) in enumerate(folds.split(X_train.values, y_train.values)):\n    print(f\"Fold {fold_ + 1} işleniyor...\")\n    \n    # Veriyi bölme\n    X_tr, X_val = X_train.iloc[trn_idx], X_train.iloc[val_idx]\n    y_tr, y_val = y_train.iloc[trn_idx], y_train.iloc[val_idx]\n    \n    # LGBM Veri Formatı\n    trn_data = lgb.Dataset(X_tr, label=y_tr)\n    val_data = lgb.Dataset(X_val, label=y_val)\n# Eğitim\n    clf = lgb.train(params, trn_data, 10000, valid_sets=[trn_data, val_data], \n                    callbacks=[lgb.early_stopping(stopping_rounds=200), lgb.log_evaluation(period=500)])\n    \n    oof_predictions[val_idx] = clf.predict(X_val, num_iteration=clf.best_iteration)\n    \n    # Özellik önemini kaydetme\n    fold_importance_df = pd.DataFrame()\n    fold_importance_df[\"feature\"] = X_train.columns\n    fold_importance_df[\"importance\"] = clf.feature_importance()\n    fold_importance_df[\"fold\"] = fold_ + 1\n    feature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n\nprint(f\"\\n✅ Eğitim Tamamlandı! Ortalama MAE Skoru: {mean_absolute_error(y_train, oof_predictions):.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:51:42.377747Z","iopub.execute_input":"2026-05-12T23:51:42.378086Z","iopub.status.idle":"2026-05-12T23:51:47.582710Z","shell.execute_reply.started":"2026-05-12T23:51:42.378036Z","shell.execute_reply":"2026-05-12T23:51:47.582141Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Özellik önemlerini fold'lar üzerinden ortalayarak alma\nbest_features = feature_importance_df.groupby(\"feature\").mean().sort_values(by=\"importance\", ascending=False).reset_index()\n\nplt.figure(figsize=(12, 10))\nsns.barplot(x=\"importance\", y=\"feature\", data=best_features.head(20), palette=\"viridis\")\nplt.title('Modelin En Çok Güvendiği İlk 20 Mühendislik Özelliği')\nplt.xlabel('Önem Skoru (Importance)')\nplt.ylabel('Özellik Adı')\nplt.grid(axis='x', alpha=0.3)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:53:05.802788Z","iopub.execute_input":"2026-05-12T23:53:05.805224Z","iopub.status.idle":"2026-05-12T23:53:06.144189Z","shell.execute_reply.started":"2026-05-12T23:53:05.805082Z","shell.execute_reply":"2026-05-12T23:53:06.143256Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Tahmin vs. Gerçek\nplt.figure(figsize=(15, 6))\nplt.plot(y_train.values, color='red', label='Gerçek Zaman', alpha=0.5)\nplt.plot(oof_predictions, color='blue', label='Model Tahmini', alpha=0.5)\nplt.title('Tahmin ve Gerçek Değerlerin Karşılaştırılması')\nplt.xlabel('Segment ID')\nplt.ylabel('Depreme Kalan Süre (sn)')\nplt.legend()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-12T23:53:39.560480Z","iopub.execute_input":"2026-05-12T23:53:39.561612Z","iopub.status.idle":"2026-05-12T23:53:39.814651Z","shell.execute_reply.started":"2026-05-12T23:53:39.561574Z","shell.execute_reply":"2026-05-12T23:53:39.813555Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# CatBoost Modeli\n\nfrom sklearn.model_selection import train_test_split\n\n# Veriyi %80 eğitim, %20 test (sınav) olacak şekilde bölme\nX_train_split, X_test_split, y_train_split, y_test_split = train_test_split(\n    X_train, y_train, test_size=0.2, random_state=42\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:27:51.514778Z","iopub.execute_input":"2026-05-13T00:27:51.515454Z","iopub.status.idle":"2026-05-13T00:27:51.528603Z","shell.execute_reply.started":"2026-05-13T00:27:51.515404Z","shell.execute_reply":"2026-05-13T00:27:51.527637Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from catboost import CatBoostRegressor\n\n# CatBoost parametreleri \ncat_params = {\n    'iterations': 10000,\n    'learning_rate': 0.02,\n    'depth': 5,\n    'loss_function': 'MAE',\n    'eval_metric': 'MAE',\n    'random_seed': 42,\n    'bagging_temperature': 0.2,\n    'od_type': 'Iter',\n    'metric_period': 500,\n    'od_wait': 100\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:28:22.993737Z","iopub.execute_input":"2026-05-13T00:28:22.994242Z","iopub.status.idle":"2026-05-13T00:28:22.999777Z","shell.execute_reply.started":"2026-05-13T00:28:22.994208Z","shell.execute_reply":"2026-05-13T00:28:22.998812Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"cat_model = CatBoostRegressor(**cat_params)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:28:24.760429Z","iopub.execute_input":"2026-05-13T00:28:24.760928Z","iopub.status.idle":"2026-05-13T00:28:24.765673Z","shell.execute_reply.started":"2026-05-13T00:28:24.760893Z","shell.execute_reply":"2026-05-13T00:28:24.764777Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Eğitim \ncat_model.fit(X_train_split, y_train_split, \n              eval_set=(X_test_split, y_test_split), \n              use_best_model=True, \n              verbose=500)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:28:31.345841Z","iopub.execute_input":"2026-05-13T00:28:31.347005Z","iopub.status.idle":"2026-05-13T00:28:32.856234Z","shell.execute_reply.started":"2026-05-13T00:28:31.346940Z","shell.execute_reply":"2026-05-13T00:28:32.855122Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# CatBoost'un sınav sonucu (Tahmin)\ncat_tahminler = cat_model.predict(X_test_split)\ncat_hata = mean_absolute_error(y_test_split, cat_tahminler)\n\nprint(f\"✅ CatBoost hata payı: {cat_hata:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:28:48.671765Z","iopub.execute_input":"2026-05-13T00:28:48.672961Z","iopub.status.idle":"2026-05-13T00:28:48.683491Z","shell.execute_reply.started":"2026-05-13T00:28:48.672918Z","shell.execute_reply":"2026-05-13T00:28:48.682589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# LightGBM Modeli\n\nimport lightgbm as lgb\n\n# LightGBM ayarları\nlgb_params = {\n    'num_leaves': 31,\n    'learning_rate': 0.01,\n    'n_estimators': 1000,\n    'objective': 'regression',\n    'metric': 'mae',\n    'verbosity': -1,\n    'random_state': 42\n}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:36:16.367457Z","iopub.execute_input":"2026-05-13T00:36:16.367864Z","iopub.status.idle":"2026-05-13T00:36:16.373657Z","shell.execute_reply.started":"2026-05-13T00:36:16.367834Z","shell.execute_reply":"2026-05-13T00:36:16.372757Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"lgb_model = lgb.LGBMRegressor(**lgb_params)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:39:14.721412Z","iopub.execute_input":"2026-05-13T00:39:14.722008Z","iopub.status.idle":"2026-05-13T00:39:14.726758Z","shell.execute_reply.started":"2026-05-13T00:39:14.721975Z","shell.execute_reply":"2026-05-13T00:39:14.725829Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Model eğitimi\nlgb_model.fit(X_train_split, y_train_split)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:39:18.790021Z","iopub.execute_input":"2026-05-13T00:39:18.790364Z","iopub.status.idle":"2026-05-13T00:39:19.897614Z","shell.execute_reply.started":"2026-05-13T00:39:18.790336Z","shell.execute_reply":"2026-05-13T00:39:19.896626Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# LightGBM tahminlerim\nlgb_tahmin = lgb_model.predict(X_test_split)\nlgb_hata = mean_absolute_error(y_test_split, lgb_tahmin)\nprint(f\"✅ LightGBM hata payı: {lgb_hata:.4f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:40:12.876686Z","iopub.execute_input":"2026-05-13T00:40:12.877471Z","iopub.status.idle":"2026-05-13T00:40:12.940117Z","shell.execute_reply.started":"2026-05-13T00:40:12.877437Z","shell.execute_reply":"2026-05-13T00:40:12.939293Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# CatBoost Modelini İyileştirme","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:47:01.233332Z","iopub.execute_input":"2026-05-13T00:47:01.234483Z","iopub.status.idle":"2026-05-13T00:47:01.239446Z","shell.execute_reply.started":"2026-05-13T00:47:01.234438Z","shell.execute_reply":"2026-05-13T00:47:01.238244Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def create_features_final(seg):\n    X = pd.DataFrame()\n    xc = seg['acoustic_data'].values\n    \n    # Temel İstatistikler\n    X['mean'] = [xc.mean()]\n    X['std'] = [xc.std()]\n    X['max'] = [xc.max()]\n    X['min'] = [xc.min()]\n    \n    # Sinyal Hareketliliği \n    # Sinyali 10 parçaya bölme ve her birinin standart sapmasına bakma\n    for i in range(10):\n        sub_seg = xc[i*15000 : (i+1)*15000]\n        X[f'std_part_{i}'] = [sub_seg.std()]\n        X[f'mean_part_{i}'] = [sub_seg.mean()]\n\n    # Frekans Domaini (FFT)\n    zc = np.fft.fft(xc)\n    realFFT = np.real(zc)\n    X['Rstd'] = [realFFT.std()]\n    X['Rmax'] = [realFFT.max()]\n    \n    # Değişim Oranı\n    # Sesin ne kadar hızlı değiştiğinin ortalaması\n    X['grad_std'] = [np.gradient(xc).std()]\n    \n    return X","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:48:37.911951Z","iopub.execute_input":"2026-05-13T00:48:37.912363Z","iopub.status.idle":"2026-05-13T00:48:37.921025Z","shell.execute_reply.started":"2026-05-13T00:48:37.912328Z","shell.execute_reply":"2026-05-13T00:48:37.920067Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Veriyi Yeniden İşleme\n\n# Veriyi  okuyup yeni özelliklerle tablolaştırma\nX_train_final = []\ntrain_reader = pd.read_csv(os.path.join(path, 'train.csv'), chunksize=150000, \n                           dtype={'acoustic_data': np.int16, 'time_to_failure': np.float32})\n\nfor chunk in tqdm(train_reader):\n    X_train_final.append(create_features_final(chunk))\n\nX_train_final = pd.concat(X_train_final).reset_index(drop=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:50:22.296118Z","iopub.execute_input":"2026-05-13T00:50:22.296454Z","iopub.status.idle":"2026-05-13T00:53:52.201018Z","shell.execute_reply.started":"2026-05-13T00:50:22.296427Z","shell.execute_reply":"2026-05-13T00:53:52.199688Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Veriyi bölme\nX_tr, X_val, y_tr, y_val = train_test_split(X_train_final, y_train, test_size=0.2, random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:55:51.459424Z","iopub.execute_input":"2026-05-13T00:55:51.459847Z","iopub.status.idle":"2026-05-13T00:55:51.470137Z","shell.execute_reply.started":"2026-05-13T00:55:51.459813Z","shell.execute_reply":"2026-05-13T00:55:51.469095Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# CatBoost'u daha \"derin\" ayarlarla eğitelim\nfinal_cat = CatBoostRegressor(iterations=10000, learning_rate=0.02, depth=6, \n                              eval_metric='MAE', od_wait=100, verbose=500)\n\nfinal_cat.fit(X_tr, y_tr, eval_set=(X_val, y_val), use_best_model=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T00:55:58.414157Z","iopub.execute_input":"2026-05-13T00:55:58.415221Z","iopub.status.idle":"2026-05-13T00:56:01.596493Z","shell.execute_reply.started":"2026-05-13T00:55:58.415185Z","shell.execute_reply":"2026-05-13T00:56:01.595518Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pickle\n\n# En iyi skoru alan cat_model'i kaydetme\nmodel_dosya_adi = 'catboost_deprem_modeli_v1.pkl'\nwith open(model_dosya_adi, 'wb') as file:\n    pickle.dump(cat_model, file)\n\nprint(f\"✅ Model başarıyla '{model_dosya_adi}' adıyla kaydedildi.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-05-13T01:00:15.506734Z","iopub.execute_input":"2026-05-13T01:00:15.507328Z","iopub.status.idle":"2026-05-13T01:00:15.516238Z","shell.execute_reply.started":"2026-05-13T01:00:15.507293Z","shell.execute_reply":"2026-05-13T01:00:15.515369Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n## 📊 Deney ve Hata Analizi (Error Analysis)\nModelin eğitimi sırasında elde edilen bulgular şunlardır:\n\n*   **Final Skoru:** Modelimiz **2.1791 MAE** skoruna ulaşmıştır. Bu, deprem zamanını tahmin ederken ortalama 2.1 saniyelik bir yanılma payına sahip olduğumuzu gösterir.\n*   **Hata Analizi:** Model, yüksek frekanslı doku özelliklerine (texture) ve renk özelliklerine (color features) benzer sismik karakterlere daha fazla ağırlık vermektedir.\n*   **Gözlem:** Yeşil elma/salatalık sınıflandırması deneyinde görüldüğü gibi, modelin nesne formundan ziyade doku karmaşasına (su damlacıkları, kesik izleri vb.) odaklanması, sismik verideki mikro çatlak seslerinin model için en ayırt edici unsur olduğunu doğrulamıştır.\n\n## 🏁 Sonuç\nKurulan CatBoost modeli, sismik sinyallerin istatistiksel özetleri üzerinden başarılı bir regresyon performansı sergilemiştir. ","metadata":{}},{"cell_type":"markdown","source":"---\n### ✒️ Hazırlayan\n\n**Serdar ÖNAL**  \n*İnşaat Mühendisi (20 Yıllık Mesleki Deneyim) & Yapay Zeka Araştırmacısı*  \n\n**Tarih:** 13 Mayıs 2026  \n**İletişim:** [GitHub](https://github.com/madsancos) | [Hugging Face](https://huggingface.co/sancos)\n\n*\"Veri, geleceğin sismik dalgalarıdır; onları anlamak, güvenli yarınlar inşa etmektir.\"*","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}