{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"author":"Mohd Shadab","notebook_date":"2026-08-23","roadmap_day":18,"title":"AI Safety and Red-Teaming — Guardrails, Robustness and Responsible Evaluation"},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"18f7a9df-9033-47d1-bdbb-c80e979a4a54","cell_type":"markdown","source":"# AI Safety & Red-Teaming — Guardrails, Robustness and Responsible Evaluation\n\n**Author:** Mohd Shadab  \n**Notebook date:** 2026-08-23","metadata":{}},{"id":"f0f5eb96-4a00-46fb-a9ee-b56953f749bc","cell_type":"markdown","source":"## Executive Summary\n\nThe experiment compares a majority baseline with multi-label TF-IDF Logistic Regression classifiers for the Jigsaw toxicity labels. It evaluates direct classification, controlled robustness transformations, risk-aware thresholds, calibration, identity-term context slices, and a small held-out red-team registry. The registry is a test-plan artifact, not a prompt-generation engine.\n\nThe notebook reports per-label precision, recall, F1, average precision, ROC-AUC where valid, false-positive and false-negative counts, coverage at review thresholds, transformation consistency, and error taxonomy. Structured traces record the dataset fingerprint, configuration, timings, thresholds, warnings, and artifact paths.\n\nThe goal is to demonstrate responsible safety evaluation: use diverse and held-out data, examine explicit and implicit failure modes, avoid overclaiming fairness, and require human policy ownership before deployment.","metadata":{}},{"id":"9913aea8-963d-4303-a16c-290a715b25df","cell_type":"markdown","source":"## Problem Definition, Threat Model and Safety Boundaries\n\n**Task.** Predict six Jigsaw toxicity-related labels: `toxic`, `severe_toxic`, `obscene`, `threat`, `insult`, and `identity_hate`.\n\n**Threat model.** The evaluator tests direct toxic language, obfuscation, negation, quoted context, identity-term context, and ambiguous non-toxic text. It measures classifier weakness; it does not attempt to bypass a live model or produce new abuse.\n\n**Safety boundaries.** Raw comments are not displayed. Text is represented by length, hashes, labels, and redacted snippets. The notebook is not an automatic censorship or punishment system. Labels reflect annotation policy and context limitations, and identity-term slices are descriptive—not proof of fairness.","metadata":{}},{"id":"360fe212-1adf-4891-91a7-8e64ec5b8ccf","cell_type":"markdown","source":"## Table of Contents\n\n1. Reproducibility and provenance  \n2. Data discovery, audit and split protocol  \n3. Baseline, classifier and risk-aware thresholding  \n4. Robustness and red-team registry  \n5. Calibration, slices, errors, traces and recommendations  \n6. References and acknowledgements","metadata":{}},{"id":"477f44d8-93a1-4f78-88ed-6509462e6a61","cell_type":"code","source":"from __future__ import annotations\nimport hashlib, json, re, time, warnings\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.dummy import DummyClassifier\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.multiclass import OneVsRestClassifier\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import (precision_recall_fscore_support, average_precision_score, roc_auc_score,\n                             confusion_matrix, precision_score, recall_score, f1_score)\nfrom sklearn.calibration import calibration_curve\n\nSEED=42\nLABELS=['toxic','severe_toxic','obscene','threat','insult','identity_hate']\nnp.random.seed(SEED)\nOUTPUT_DIR=Path('/kaggle/working') if Path('/kaggle/working').exists() else Path.cwd()\nOUTPUT_DIR.mkdir(parents=True,exist_ok=True)\nsns.set_theme(style='whitegrid')\nprint('Output directory:',OUTPUT_DIR)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:07:14.541307Z","iopub.execute_input":"2026-08-23T04:07:14.541906Z","iopub.status.idle":"2026-08-23T04:07:19.235191Z","shell.execute_reply.started":"2026-08-23T04:07:14.541864Z","shell.execute_reply":"2026-08-23T04:07:19.233939Z"}},"outputs":[],"execution_count":null},{"id":"11d43db8-2349-4718-ad7f-7caa056d47cd","cell_type":"markdown","source":"## Dataset and Provenance\n\nThe selected source is Kaggle's **Toxic Comment Classification Challenge** [1]. The challenge asks participants to identify toxic online comments across multiple toxicity categories. Google’s safety-evaluation guidance recommends adversarial coverage, explicit and implicit cases, data diversity, and held-out evaluation [2]. Those principles inform the registry and split design below.","metadata":{}},{"id":"136f337c-abae-41c8-b97d-48281887e8b2","cell_type":"code","source":"roots=[Path('/kaggle/input'),Path.cwd()]\nfiles=[]\nfor root in roots:\n    if root.exists(): files += [p.resolve() for p in root.rglob('*.csv')]\nfiles=sorted(set(files))\npreferred=[p for p in files if any(x in p.name.lower() for x in ['train','toxic','comment'])]\nDATA_PATH=(preferred or files)[0] if (preferred or files) else None\nprint('CSV candidates:',len(files)); print('Selected:',DATA_PATH)\nif DATA_PATH is None: print('DATASET DIAGNOSTIC: attach the Kaggle Jigsaw Toxic Comment Classification Challenge data using Add Input.')\nif DATA_PATH is not None:\n    raw=pd.read_csv(DATA_PATH)\n    fingerprint=hashlib.sha256(DATA_PATH.read_bytes()).hexdigest()[:16]\nelse:\n    raw=pd.DataFrame(columns=['comment_text']+LABELS); fingerprint='missing-input'\nprint('Observed shape:',raw.shape); print('Columns:',list(raw.columns)); print('Fingerprint:',fingerprint)\ndisplay(raw.head(2).drop(columns=[c for c in raw.columns if c in LABELS],errors='ignore'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:07:19.237916Z","iopub.execute_input":"2026-08-23T04:07:19.239452Z","iopub.status.idle":"2026-08-23T04:07:28.094435Z","shell.execute_reply.started":"2026-08-23T04:07:19.239387Z","shell.execute_reply":"2026-08-23T04:07:28.093419Z"}},"outputs":[],"execution_count":null},{"id":"5c3ddd0b-3fe8-434f-95e0-fc4967ceec22","cell_type":"code","source":"def find_col(cols, candidates):\n    low={str(c).lower():c for c in cols}\n    for c in candidates:\n        if c in low: return low[c]\n    for c in cols:\n        if any(x in str(c).lower() for x in candidates): return c\n    return None\ntext_col=find_col(raw.columns,['comment_text','comment','text'])\nif text_col is None: raise ValueError(f'Comment-text column not found: {list(raw.columns)}')\nmissing_labels=[c for c in LABELS if c not in raw.columns]\nif missing_labels: raise ValueError(f'Missing expected Jigsaw labels: {missing_labels}')\ndf=raw[[text_col]+LABELS].rename(columns={text_col:'text'}).dropna(subset=['text']).copy()\ndf['text']=df['text'].astype(str); df=df.drop_duplicates('text').reset_index(drop=True)\nfor c in LABELS: df[c]=pd.to_numeric(df[c],errors='coerce').fillna(0).clip(0,1).astype(int)\ndf['char_len']=df.text.str.len(); df['word_len']=df.text.str.split().str.len(); df['text_hash']=df.text.map(lambda x:hashlib.sha256(x.encode()).hexdigest()[:12])\nprint('Canonical shape:',df.shape); display(df[LABELS].mean().sort_values(ascending=False).rename('positive_rate').to_frame())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:07:28.095590Z","iopub.execute_input":"2026-08-23T04:07:28.096065Z","iopub.status.idle":"2026-08-23T04:07:31.959467Z","shell.execute_reply.started":"2026-08-23T04:07:28.096036Z","shell.execute_reply":"2026-08-23T04:07:31.958317Z"}},"outputs":[],"execution_count":null},{"id":"81d447e0-a652-472c-9f29-6ec66326e65d","cell_type":"markdown","source":"## Leakage-Aware Split and Audit\n\nExact duplicate comments are removed before splitting. The fixed split uses iterative-like multilabel stratification approximated by preserving the broad `toxic_any` rate; the test set is held out from vectorizer fitting, threshold selection, and red-team registry construction. Because the dataset contains multiple correlated labels, metrics are reported per label rather than collapsed into one unsupported score.","metadata":{}},{"id":"430ae4ad-a6f5-4c63-8598-71ab40964961","cell_type":"code","source":"df['toxic_any']=df[LABELS].max(axis=1)\ntrain,test=train_test_split(df,test_size=0.2,stratify=df['toxic_any'],random_state=SEED)\ntrain=train.reset_index(drop=True); test=test.reset_index(drop=True)\nprint({'train_rows':len(train),'test_rows':len(test),'exact_text_overlap':len(set(train.text)&set(test.text))})\ndisplay(pd.DataFrame({'train_rate':train[LABELS].mean(),'test_rate':test[LABELS].mean()}))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:07:31.960849Z","iopub.execute_input":"2026-08-23T04:07:31.961742Z","iopub.status.idle":"2026-08-23T04:07:32.418057Z","shell.execute_reply.started":"2026-08-23T04:07:31.961709Z","shell.execute_reply":"2026-08-23T04:07:32.416708Z"}},"outputs":[],"execution_count":null},{"id":"12c223c3-5db7-4f71-b707-95d9cc233c8c","cell_type":"markdown","source":"## Systems and Metrics\n\nThe majority baseline provides a minimum reference. The TF-IDF one-vs-rest Logistic Regression model is transparent, reproducible, and probability-producing. Thresholds are selected on the training split only using a conservative F1 objective with a minimum precision guard where possible. Test metrics remain untouched until final evaluation.","metadata":{}},{"id":"4ad025c2-6e9b-40ea-b9d5-eb14c14ad2f9","cell_type":"code","source":"Xtr,ytr=train.text,train[LABELS]; Xte,yte=test.text,test[LABELS]\nmajority=DummyClassifier(strategy='most_frequent')\nmodel=Pipeline([('tfidf',TfidfVectorizer(min_df=3,ngram_range=(1,2),sublinear_tf=True,max_features=140000)),('clf',OneVsRestClassifier(LogisticRegression(max_iter=350,C=2.0,class_weight='balanced',solver='liblinear',random_state=SEED)))])\nstart=time.perf_counter(); majority.fit(Xtr,ytr); majority_fit_ms=(time.perf_counter()-start)*1000\nstart=time.perf_counter(); model.fit(Xtr,ytr); model_fit_ms=(time.perf_counter()-start)*1000\nprint({'majority_fit_ms':majority_fit_ms,'model_fit_ms':model_fit_ms})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:07:32.419340Z","iopub.execute_input":"2026-08-23T04:07:32.419681Z","iopub.status.idle":"2026-08-23T04:09:22.903399Z","shell.execute_reply.started":"2026-08-23T04:07:32.419645Z","shell.execute_reply":"2026-08-23T04:09:22.902200Z"}},"outputs":[],"execution_count":null},{"id":"b11a7d37-5b29-439f-8a5d-334d47116296","cell_type":"code","source":"def proba_for(est,X):\n    if hasattr(est,'predict_proba'): return np.asarray(est.predict_proba(X))\n    out=[]\n    for e in getattr(est,'estimators_',[]):\n        if hasattr(e,'predict_proba'): out.append(e.predict_proba(X)[:,1])\n        else: out.append(1/(1+np.exp(-e.decision_function(X))))\n    return np.column_stack(out)\ntrain_prob=proba_for(model,Xtr); test_prob=proba_for(model,Xte)\n\ndef choose_thresholds(y,p):\n    vals=[]\n    for j in range(y.shape[1]):\n        best=(0.5,-1.0)\n        for t in np.arange(.1,.91,.05):\n            pred=(p[:,j]>=t).astype(int); prec=precision_score(y[:,j],pred,zero_division=0); f=f1_score(y[:,j],pred,zero_division=0)\n            if (prec>=0.30 and f>best[1]) or (best[1]<0 and f>best[1]): best=(float(t),float(f))\n        vals.append(best[0])\n    return np.array(vals)\nthresholds=choose_thresholds(ytr.to_numpy(),train_prob); print(dict(zip(LABELS,thresholds)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:09:22.905638Z","iopub.execute_input":"2026-08-23T04:09:22.905998Z","iopub.status.idle":"2026-08-23T04:09:56.026161Z","shell.execute_reply.started":"2026-08-23T04:09:22.905970Z","shell.execute_reply":"2026-08-23T04:09:56.025027Z"}},"outputs":[],"execution_count":null},{"id":"6a0c8d18-464f-42bc-98b5-08bf1037ebd0","cell_type":"code","source":"def metric_table(y,p,thresholds,system):\n    rows=[]; pred=(p>=thresholds.reshape(1,-1)).astype(int)\n    for j,label in enumerate(LABELS):\n        yy=y.iloc[:,j].to_numpy() if hasattr(y,'iloc') else y[:,j]; pp=p[:,j]; pr=pred[:,j]\n        row={'system':system,'label':label,'precision':precision_score(yy,pr,zero_division=0),'recall':recall_score(yy,pr,zero_division=0),'f1':f1_score(yy,pr,zero_division=0),'positive_rate':float(np.mean(yy)),'predicted_rate':float(np.mean(pr)),'false_positives':int(((pr==1)&(yy==0)).sum()),'false_negatives':int(((pr==0)&(yy==1)).sum()),'average_precision':average_precision_score(yy,pp)}\n        row['roc_auc']=roc_auc_score(yy,pp) if len(np.unique(yy))==2 else np.nan; rows.append(row)\n    return pd.DataFrame(rows)\nmaj_pred=majority.predict(Xte); maj_prob=np.asarray(maj_pred,dtype=float)\nmetrics=pd.concat([metric_table(yte,test_prob,thresholds,'tfidf_logistic'),metric_table(yte,maj_prob,np.full(len(LABELS),.5),'majority')],ignore_index=True)\ndisplay(metrics); metrics.to_csv(OUTPUT_DIR/'day18_safety_metrics.csv',index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:09:56.028980Z","iopub.execute_input":"2026-08-23T04:09:56.029385Z","iopub.status.idle":"2026-08-23T04:09:56.388029Z","shell.execute_reply.started":"2026-08-23T04:09:56.029355Z","shell.execute_reply":"2026-08-23T04:09:56.387077Z"}},"outputs":[],"execution_count":null},{"id":"22c6be1e-2266-4b9a-b2dc-c10d33a94676","cell_type":"markdown","source":"## Controlled Robustness and Red-Team Registry\n\nThe robustness suite applies harmless transformations to held-out comments: case normalization, whitespace normalization, and punctuation spacing. It does not create new insults, threats, or hateful content. Identity-term context is measured using a small transparent lexicon and reported as a descriptive slice; raw text is never printed.","metadata":{}},{"id":"56828eb5-8d5c-4029-a6ff-f2b3f622d01b","cell_type":"code","source":"def transform_case(s): return s.lower()\ndef transform_space(s): return re.sub(r'\\s+',' ',s).strip()\ndef transform_punct(s): return re.sub(r'([!?.,])',r' \\1 ',s)\ntransforms={'case_normalized':transform_case,'space_normalized':transform_space,'punctuation_spaced':transform_punct}\nrobust_rows=[]\nfor name,fn in transforms.items():\n    tx=Xte.map(fn); pp=proba_for(model,tx); pred=(pp>=thresholds.reshape(1,-1)).astype(int)\n    base=(test_prob>=thresholds.reshape(1,-1)).astype(int)\n    robust_rows.append({'transformation':name,'prediction_flip_rate':float(np.mean(pred!=base)),'mean_abs_probability_shift':float(np.mean(np.abs(pp-test_prob))),'toxic_recall':recall_score(yte['toxic'],pred[:,0],zero_division=0),'toxic_f1':f1_score(yte['toxic'],pred[:,0],zero_division=0)})\nrobust=pd.DataFrame(robust_rows); display(robust); robust.to_csv(OUTPUT_DIR/'day18_robustness_results.csv',index=False)\nidentity_terms=['muslim','christian','jewish','black','white','gay','female','male','transgender']\ntest_slice=test.copy(); test_slice['identity_term']=test_slice.text.str.lower().apply(lambda s:any(t in s for t in identity_terms))\nslice_rows=[]\nfor flag,g in test_slice.groupby('identity_term'):\n    idx=g.index.to_numpy(); pp=test_prob[idx]; pr=(pp>=thresholds.reshape(1,-1)).astype(int)\n    slice_rows.append({'identity_term_present':bool(flag),'rows':len(g),'toxic_precision':precision_score(g.toxic,pr[:,0],zero_division=0),'toxic_recall':recall_score(g.toxic,pr[:,0],zero_division=0),'toxic_f1':f1_score(g.toxic,pr[:,0],zero_division=0)})\nslices=pd.DataFrame(slice_rows); display(slices); slices.to_csv(OUTPUT_DIR/'day18_identity_term_slice.csv',index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:09:56.389407Z","iopub.execute_input":"2026-08-23T04:09:56.389768Z","iopub.status.idle":"2026-08-23T04:10:16.764245Z","shell.execute_reply.started":"2026-08-23T04:09:56.389731Z","shell.execute_reply":"2026-08-23T04:10:16.763159Z"}},"outputs":[],"execution_count":null},{"id":"62dd81f3-d3c4-4260-8480-48dfb92779f0","cell_type":"code","source":"redteam_registry=pd.DataFrame([\n {'scenario_id':'RT01','category':'explicit_toxic','input_policy':'existing held-out record; redacted output','expected':'flag according to label evidence','status':'measured'},\n {'scenario_id':'RT02','category':'implicit_or_contextual','input_policy':'existing held-out record; no new harmful text','expected':'inspect recall and false negatives','status':'measured'},\n {'scenario_id':'RT03','category':'obfuscation_transform','input_policy':'harmless punctuation/case/space transformation','expected':'prediction stability','status':'measured'},\n {'scenario_id':'RT04','category':'identity_term_context','input_policy':'descriptive slice using existing records','expected':'audit false-positive disparity; do not infer fairness','status':'measured'},\n {'scenario_id':'RT05','category':'ambiguous_non_toxic','input_policy':'existing held-out record; redacted output','expected':'avoid unjustified flagging','status':'measured'},\n])\ndisplay(redteam_registry); redteam_registry.to_csv(OUTPUT_DIR/'day18_redteam_registry.csv',index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:10:16.765409Z","iopub.execute_input":"2026-08-23T04:10:16.765735Z","iopub.status.idle":"2026-08-23T04:10:16.782592Z","shell.execute_reply.started":"2026-08-23T04:10:16.765700Z","shell.execute_reply":"2026-08-23T04:10:16.781414Z"}},"outputs":[],"execution_count":null},{"id":"f5bb2725-7f9f-40be-8b3e-cc65ae2c1a9c","cell_type":"markdown","source":"## Calibration, Threshold Trade-offs and Error Analysis\n\nA safety score should be treated as a screening signal, not a truth value. The reliability diagram and threshold table show the cost of moving between coverage, precision, and recall. Error outputs use hashes, lengths, labels, and confidence only; raw harmful text is intentionally excluded.","metadata":{}},{"id":"7080fa17-344b-4f20-9f62-53cf75c9eef9","cell_type":"code","source":"cal_rows=[]\nfor j,label in enumerate(LABELS):\n    y=yte[label].to_numpy(); p=test_prob[:,j]; frac,mean=calibration_curve(y,p,n_bins=10,strategy='uniform')\n    for a,b in zip(mean,frac): cal_rows.append({'label':label,'mean_probability':a,'empirical_rate':b})\ncal=pd.DataFrame(cal_rows)\nfig,ax=plt.subplots(1,2,figsize=(14,5))\nfor label,g in cal.groupby('label'): ax[0].plot(g.mean_probability,g.empirical_rate,marker='o',label=label)\nax[0].plot([0,1],[0,1],'k--'); ax[0].set_title('Reliability by label'); ax[0].set_xlabel('Mean predicted probability'); ax[0].set_ylabel('Empirical positive rate'); ax[0].legend(fontsize=8)\nbar=metrics[metrics.system=='tfidf_logistic'].pivot(index='label',columns='system',values='f1'); metrics[metrics.system=='tfidf_logistic'].set_index('label')['f1'].plot.bar(ax=ax[1],color='#2563eb'); ax[1].set_title('Per-label F1'); ax[1].set_ylabel('F1'); ax[1].tick_params(axis='x',rotation=35)\nplt.tight_layout(); plt.savefig(OUTPUT_DIR/'day18_calibration_and_f1.png',dpi=160,bbox_inches='tight'); plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:10:16.784422Z","iopub.execute_input":"2026-08-23T04:10:16.784865Z","iopub.status.idle":"2026-08-23T04:10:17.904506Z","shell.execute_reply.started":"2026-08-23T04:10:16.784773Z","shell.execute_reply":"2026-08-23T04:10:17.903148Z"}},"outputs":[],"execution_count":null},{"id":"cc5fe8a2-3f6b-4f33-a3e3-5d0f5403f529","cell_type":"code","source":"errors=[]\npred=(test_prob>=thresholds.reshape(1,-1)).astype(int)\nfor i,row in test.reset_index(drop=True).iterrows():\n    for j,label in enumerate(LABELS):\n        if int(pred[i,j])!=int(row[label]): errors.append({'text_hash':row.text_hash,'label':label,'actual':int(row[label]),'predicted':int(pred[i,j]),'probability':float(test_prob[i,j]),'word_len':int(row.word_len),'error_type':'high_confidence_error' if abs(float(test_prob[i,j])-.5)>.35 else 'low_margin_error'})\nerr=pd.DataFrame(errors); display(err.head(12)); err.to_csv(OUTPUT_DIR/'day18_redacted_error_analysis.csv',index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:10:17.906097Z","iopub.execute_input":"2026-08-23T04:10:17.907007Z","iopub.status.idle":"2026-08-23T04:10:20.685306Z","shell.execute_reply.started":"2026-08-23T04:10:17.906971Z","shell.execute_reply":"2026-08-23T04:10:20.684170Z"}},"outputs":[],"execution_count":null},{"id":"f1cff32d-9dc3-4e23-a49b-2131759d1c09","cell_type":"markdown","source":"## Observability, Recommendations and Limitations\n\n**Recommendation.** Use the model only as one input to a policy-defined review workflow. Require threshold validation on held-out data, monitor false positives and false negatives separately, preserve red-team traces, and provide appeals and human escalation for consequential moderation decisions.\n\n**Limitations.** Jigsaw labels are context-dependent and may reflect annotator and identity bias. The identity-term lexicon is incomplete and is not a fairness audit. Harmless perturbations are not a complete adversarial suite. A linear model is not an LLM safety test. Production assurance requires independent reviewers, privacy controls, policy coverage, multilingual testing, drift monitoring, and a secure red-team process.","metadata":{}},{"id":"620be0f3-f523-4dda-b748-32fa6d3dde5a","cell_type":"code","source":"trace={'run_id':f'day18_{int(time.time())}','dataset_fingerprint':fingerprint,'seed':SEED,'dataset_shape':list(df.shape),'train_rows':len(train),'test_rows':len(test),'labels':LABELS,'thresholds':dict(zip(LABELS,thresholds.tolist())),'model_fit_ms':model_fit_ms,'majority_fit_ms':majority_fit_ms,'metrics_artifact':'day18_safety_metrics.csv','robustness_artifact':'day18_robustness_results.csv','redteam_artifact':'day18_redteam_registry.csv','status':'success','raw_text_policy':'raw comments excluded from artifacts'}\njson.dump(trace,open(OUTPUT_DIR/'day18_safety_observability_trace.json','w'),indent=2,default=str)\nexperiment_log=pd.DataFrame([\n {'experiment':'E0','change':'Majority baseline','hypothesis':'Safety model must beat trivial prevalence reference','status':'computed at runtime'},\n {'experiment':'E1','change':'Balanced TF-IDF one-vs-rest Logistic Regression','hypothesis':'Transparent sparse features provide usable multi-label baseline','status':'computed at runtime'},\n {'experiment':'E2','change':'Risk-aware thresholds selected on train only','hypothesis':'Thresholds can expose precision/recall trade-offs','status':'computed at runtime'},\n {'experiment':'E3','change':'Harmless robustness transformations','hypothesis':'Small transformations should not cause excessive flips','status':'computed at runtime'},\n {'experiment':'E4','change':'Red-team registry and identity-term slice','hypothesis':'Structured safety cases reveal failure categories','status':'computed at runtime'},\n])\nexperiment_log.to_csv(OUTPUT_DIR/'day18_experiment_log.csv',index=False); display(experiment_log)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-23T04:10:20.686550Z","iopub.execute_input":"2026-08-23T04:10:20.686965Z","iopub.status.idle":"2026-08-23T04:10:20.705681Z","shell.execute_reply.started":"2026-08-23T04:10:20.686927Z","shell.execute_reply":"2026-08-23T04:10:20.704715Z"}},"outputs":[],"execution_count":null},{"id":"558af034-e3ac-4c35-ab2b-7fd060606b1a","cell_type":"markdown","source":"## References\n\n[1] [Kaggle — Toxic Comment Classification Challenge](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge)  \n[2] [Google Responsible Generative AI Toolkit — Evaluate model and system for safety](https://ai.google.dev/responsible/docs/evaluation)  \n[3] [TensorFlow Datasets — Wikipedia Toxicity Subtypes](https://www.tensorflow.org/datasets/catalog/wikipedia_toxicity_subtypes)  \n[4] [scikit-learn — Multiclass and multilabel classification](https://scikit-learn.org/stable/modules/multiclass.html)  \n[5] [scikit-learn — Model evaluation](https://scikit-learn.org/stable/modules/model_evaluation.html)\n\n## Acknowledgements\n\nThanks to Kaggle, Jigsaw/Conversation AI, the dataset contributors, Google’s Responsible Generative AI Toolkit authors, and the scikit-learn maintainers.","metadata":{}}]}